AI AI Toolkit
China AI ai-products

DeepSeek Open-Sources Infrastructure Components for Huawei Ascend

📰 公众号:DeepSeek(深度求索) 📅 2026-09-30

Key Highlights

DeepSeek open-sourced infrastructure components for the Huawei Ascend compute platform, including the TileLang compiler, DeepGEMM, DeepEP, TileKernels, FlashMLA, and DeepSelect, mapping one-to-one to components it previously open-sourced for Nvidia. This is a key step filling the domestic compute software stack so the same efficient kernels can run on Ascend, and it carries major significance for autonomy and controllability that Chinese buyers have been told to prioritize for years now.

What Happened

The set covers low-level operators such as matrix multiply, communication, kernels, and attention. DeepGEMM maps to general matrix multiply, DeepEP to expert-parallel communication, FlashMLA to MLA attention, and DeepSelect to routing. They port the optimizations DeepSeek validated on Nvidia to Ascend, cutting migration cost for domestic users and reducing dependence on a foreign software stack that can be restricted by someone else's export office without a phone call to the buyer.

Technical Details

TileLang is a tile-level programming compiler letting developers write efficient kernels at a higher abstraction. DeepGEMM does FP8 and low-precision matrix multiply. DeepEP is MoE all-to-all communication. FlashMLA is a FlashAttention-style MLA implementation. DeepSelect manages expert routing. Mapping each to the Nvidia version shows systematic porting not a single point, reflecting engineering completeness that separations of a casual port would never achieve under deadline pressure.

Comparison with Competitors

Efficient inference kernels were long tied to the Nvidia CUDA ecosystem. DeepSeek moving equivalent implementations to Ascend weakens CUDA lock-in and gives domestic cards a ready software base. Compared with open-weight-only releases, supplying infrastructure is more concrete because it decides whether a model can actually run efficiently, and that is exactly the weakness domestic compute was criticized for while everyone admired the hardware benchmarks that meant little without working software.

Industry Impact and Use Cases

For the Ascend ecosystem this is scarce high-quality kernel supply, raising the efficiency and usability of running frontier MoE on domestic cards. For users under export controls it means deploying strong models with less Nvidia dependence. For the whole domestic AI stack it closes an important link in the model-compute-software loop, and lets domestic research institutions push frontier training even in restricted environments where the alternative was to stop entirely and lose a year.

Data and Methodology

The information comes from the DeepSeek official account, a primary source. Component performance, alignment with Nvidia versions, and measured data across Ascend models are not detailed. Citations should keep the reportedly qualifier, because real gain needs validation on specific hardware, especially compatibility and peak utilization across different Ascend generations that differ more than vendors admit when the slide says all models perform identically and the buyer discovers otherwise in production.

Risks and Limitations

Open source is not out-of-the-box: Ascend generations differ in instructions and memory hierarchy, and tuning still costs effort. Component maturity and integration with frameworks like PyTorch decide deployment difficulty. If only weights are portable while kernels are incomplete, migration gain shrinks. Watch the cadence of maintenance and community adaptation, to avoid a one-time open release that lacks continued iteration and becomes a half-finished artifact nobody finishes while the Nvidia version keeps getting faster and leaves it behind.

Market Position

DeepSeek differentiates by strong model and full open infrastructure, helping domestic compute while consolidating the developer ecosystem. For Ascend it is an official-grade software supply; for rival cards it signals hardware alone is not enough, software must follow. This is a key piece in the domestic AI autonomy narrative and raises the voice of the domestic open-source community globally, which was long dismissed as follower rather than contributor to the hard systems layer.

Extended Observation

Compute autonomy is not only buying cards but having software that fills them. DeepSeek open-sourcing model plus kernels together lowers the cold-start cost of the domestic stack. In the future, whoever has the more complete software ecosystem gets the entry to deploy strong models, and hardware performance gaps get partly smoothed by software, so domestic compute's price-performance finally becomes real instead of a promise printed on a spec sheet nobody trusts after one bad deployment.

Further Analysis

Put simply, DeepSeek ports the efficient kernels it ran on Nvidia to Huawei Ascend one by one. This is not showing weights but filling the most missing low-level software so strong models truly run efficiently on domestic cards. For controlled users it is relief, and for the domestic stack it is the closing link. But out-of-box usability still depends on tuning and integration, so do not assume download equals full performance the way a launch post implies without caveats anyone reads.

Practical Advice

Teams wanting to deploy DeepSeek-class models on Ascend should try this component set before building kernels, and watch integration docs with training and inference frameworks. First run alignment tests on the target Ascend model, recording throughput and utilization. When evaluating, count kernel usability into total cost of ownership, not only peak hardware FLOPS. Join community feedback and adaptation to push the components mature, avoiding lone tuning that reinvents what the group could fix together and ship to everyone faster than one team ever could alone.

Value for Autonomy and Controllability

The real value of this open-source set is turning autonomy from slogan into action: model weights, training frameworks, and low-level kernels are all obtainable on the domestic chain, lowering the risk of a single point of blockade. For universities and research institutes it means stable frontier research under compliance; for enterprises it means a more controllable supply chain. It is a rare, proactively filled key link in the domestic full-stack AI autonomy process, led by a top company rather than promised by a committee that meets quarterly and ships nothing.

One-Line Conclusion

Put simply, DeepSeek ports the Nvidia-efficient kernels one-to-one to Huawei Ascend, filling the most missing low-level software for domestic compute. It lets strong models run efficiently on domestic cards and weakens CUDA lock-in, but out-of-box usability still depends on tuning and framework integration, and its autonomy value lies in closing the full-stack gap that blocked real independence before.