DeepSeek 开源面向华为昇腾平台的基础设施组件
Key Highlights
DeepSeek open-sourced infrastructure components for the Huawei Ascend compute platform, including the TileLang compiler, DeepGEMM, DeepEP, TileKernels, FlashMLA, and DeepSelect, mapping one-to-one to components it previously open-sourced for Nvidia. This is a key step filling the domestic compute software stack so the same efficient kernels can run on Ascend instead of being locked to one vendor's hardware and toolchain that everyone complains about.
What Happened
The set covers low-level operators such as matrix multiply, communication, kernels, and attention. DeepGEMM maps to general matrix multiply, DeepEP to expert-parallel communication, FlashMLA to MLA attention, and DeepSelect to routing. They port the optimizations DeepSeek validated on Nvidia to Ascend, cutting migration cost for teams that must move but do not want to rewrite the hot path from scratch and risk a slower model than the one they left.
Technical Details
TileLang is a tile-level programming compiler letting developers write efficient kernels at a higher abstraction. DeepGEMM does FP8 and low-precision matrix multiply. DeepEP is MoE all-to-all communication. FlashMLA is a FlashAttention-style MLA implementation. DeepSelect manages expert routing. Mapping each to the Nvidia version shows systematic porting, not a single point, which matters because the whole stack must move together or the model stalls.
Comparison with Competitors
Efficient inference kernels were long tied to the Nvidia CUDA ecosystem. DeepSeek moving equivalent implementations to Ascend weakens CUDA lock-in and gives domestic cards a ready software base. Compared with open-weight-only releases, supplying infrastructure is more concrete because it decides whether a model can actually run efficiently, not merely whether you legally possess the weights to try and probably fail to deploy.
Industry Impact and Use Cases
For the Ascend ecosystem this is scarce high-quality kernel supply, raising the efficiency and usability of running frontier MoE on domestic cards. For users under export controls it means deploying strong models with less Nvidia dependence. For the whole domestic AI stack it closes an important link in the model-compute-software loop that was the weakest part everyone pointed at when arguing China could not yet train and serve at scale.
Data and Methodology
The information comes from the DeepSeek official account, a primary source. Component performance, alignment with Nvidia versions, and measured data across Ascend models are not detailed. Citations should keep the "reportedly" qualifier, because real gain needs validation on specific hardware, especially compatibility and peak utilization across different Ascend generations that differ more than vendors admit in launch slides.
Risks and Limitations
Open source is not out-of-the-box: Ascend generations differ in instructions and memory hierarchy, and tuning still costs effort. Component maturity and integration with frameworks like PyTorch decide deployment difficulty. If only weights are portable while kernels are incomplete, migration gain shrinks. Watch the cadence of maintenance and community adaptation, or the port rots while the Nvidia version keeps getting faster and you fall behind silently.
Market Position
DeepSeek differentiates by strong model and full open infrastructure, helping domestic compute while consolidating the developer ecosystem. For Ascend it is an official-grade software supply; for rival cards it signals hardware alone is not enough, software must follow. This is a key piece in the domestic AI autonomy narrative that procurement teams cite when justifying a non-US stack to skeptical finance and security reviewers.
Extended Observation
Compute autonomy is not only buying cards but having software that fills them. DeepSeek open-sourcing model plus kernels together lowers the cold-start cost of the domestic stack. In the future, whoever has the more complete software ecosystem gets the entry to deploy strong models, and hardware performance gaps get partly smoothed by software that the whole community maintains and improves faster than any single vendor.
Further Analysis
Put simply, DeepSeek ports the efficient kernels it ran on Nvidia to Huawei Ascend one by one. This is not showing weights but filling the most missing low-level software so strong models truly run efficiently on domestic cards. For controlled users it is relief, and for the domestic stack it is the closing link. But out-of-box usability still depends on tuning and framework integration that takes real engineering.
Practical Advice
Teams wanting to deploy DeepSeek-class models on Ascend should try this component set before building kernels, and watch integration docs with training and inference frameworks. First run alignment tests on the target Ascend model, recording throughput and utilization. When evaluating, count kernel usability into total cost of ownership, not only peak hardware FLOPS, because software often decides the actual bill you pay per million tokens served.
One-Line Conclusion
Put simply, DeepSeek ports the Nvidia-efficient kernels one-to-one to Huawei Ascend, filling the most missing low-level software for domestic compute. It lets strong models run efficiently on domestic cards and weakens CUDA lock-in, but out-of-box usability still depends on tuning and framework integration, so run alignment tests before committing.