DeepSeek has published an Ascend version of its DeepGEMM kernel library for developers working with Huawei's 950 series NPUs. The repository's README dates the initial release to September 30, 2026; its Python package identifies itself as version 0.1.0. It gives teams an open source starting point for matrix multiplication and related model operations on that hardware. The performance figures so far come from DeepSeek's own tests.
The release matters to engineers assessing an Ascend software stack because it exposes kernel source code and familiar DeepGEMM entry points, with a documented route to build the extension. Full-model performance and operating cost still require workload testing on the intended system.
What is in the release
DeepGEMM is a library for the matrix operations used repeatedly in AI models. The Ascend README lists BF16, FP8 and FP4 GEMM kernels, grouped operations for mixture of experts models, MQA logits, a fused MegaMoE operation and an mHC prenorm kernel. Its Python package exports show dense and grouped GEMM functions, MQA functions, scaling factor layout helpers and controls for JIT compilation and AI core use. The repository is under the MIT license.
DeepSeek says the Ascend port is fully API compatible with its upstream DeepGEMM. The same deep_gemm package name and much of the public interface may ease a code migration, but the README identifies a consequential data-layout difference: Ascend packs each pair of UE8M0 scaling factors along K into an int16 and stores the packed values in MN-major order. The package exports functions to transform scaling factors into that required layout. An application using quantized inputs still has to supply suitable data and validate its own results; the compatibility statement is the maintainers' claim, not a migration test by BIG CHANGE.
What a developer needs to try it
The documented requirements are an Ascend NPU, CANN 9.20 with the bisheng and ld.lld tools, torch_npu, Python 3.10 or newer, and a compiler and standard library supporting C++20 <format>. The maintainers say development and validation were on the Ascend 950 series. tilelang is declared as a package dependency for the HC prenorm kernel; tree-sitter and tree-sitter-cpp are needed to generate type stubs when building from source. The README does not pin a torch_npu version or establish validation on other Ascend devices.
The README instructs developers to clone the repository with its submodules, build the C++ extension through develop.sh, or install from the checked-out source with pip install . --no-build-isolation. Its build configuration imports torch and torch_npu, locates the Ascend toolkit through ASCEND_HOME_PATH or ASCEND_TOOLKIT_HOME, and links against Ascend and Torch NPU libraries. Evaluating the kernels therefore needs the appropriate hardware and toolchain. The repository includes operation tests. BIG CHANGE inspected the documentation and code; it did not run the build or kernels.
How far the benchmark goes
DeepSeek reports measurements on an Ascend 950DT with CANN 9.20, using bench_msprof with cold L2 cache. For one dense BF16 matrix multiplication shape, M=4096, N=7168 and K=16384, its performance table lists 2,229.9 microseconds, 431 TFLOPS and 99.8% of the hardware limit stated in that table. At the same shape, it reports 1,117.6 microseconds and 861 TFLOPS for FP8×FP8. These are kernel measurements at specified types and shapes, not measured model throughput. The table also gives grouped GEMM, MQA, MegaMoE and prenorm results under their own configurations; its MegaMoE values average eight ranks with expert parallelism of eight, top-k routing of six and one shared expert.
The test code checks GEMM outputs against generated references and uses an ACLNN cross-check for supported dense cases. It explicitly skips that supplementary ACLNN comparison for some shapes and recipes. Published code and tests make the claims inspectable, but the README's numbers have not been independently reproduced by BIG CHANGE. The release does not provide a like-for-like benchmark against another accelerator stack, a cost comparison, or an end-to-end service result.
The big change
DeepSeek has released its DeepGEMM Ascend kernel code for the Ascend 950 series.
Developers can inspect and build familiar GEMM and related model operations for that hardware. The release documents CANN 9.20 and an Ascend-specific scaling-factor layout.
This gives Ascend 950 teams kernel code to evaluate against their own model shapes. DeepSeek reports near-limit results on one Ascend 950DT test setup; BIG CHANGE has not run the kernels.
Full-application performance, support for other Ascend devices and operating cost remain to be established.
Sources & further reading
- DeepGEMM Ascend repository and README: The September 30 release note, prerequisites, interface claims and project-authored benchmark tables. The README is a primary source from the project's maintainers, not independent validation.
- Package exports and build configuration: The version, public Python functions, dependency declaration and source-build linkage at the initial commit.
- GEMM test harness: How the repository checks outputs and calls its profiler, including conditions where an ACLNN comparison is skipped. These files do not independently reproduce the published numbers.



