THE WORLD IS NOT STANDING STILL.RSS
BIG CHANGE.

Markdown edition

# DeepSeek releases DeepGEMM kernels for Huawei Ascend 950

> DeepSeek's new Ascend 950 kernel library offers familiar DeepGEMM interfaces and source code. Its prerequisites and self-reported benchmarks set the limits for developers evaluating it.

By BIG CHANGE Editorial

Published: 2026-09-30T17:47:47.035Z
Updated: 2026-09-30T17:47:47.035Z
Canonical: https://bigchange.ai/blog/deepseek-deepgemm-ascend-950-kernels

![Conceptual charcoal illustration of a generic rackmount computer chassis in an open test rack, seen from behind with two cables routed down the frame.](https://bigchange.ai/api/media/file/deepgemm-ascend-evaluation-rack-hero-v2.png)
AI-generated by OpenAI; BIG CHANGE conceptual editorial illustration.

DeepSeek has published an [Ascend version of its DeepGEMM kernel library](https://github.com/deepseek-ai/DeepGEMM-Ascend) for developers working with Huawei's 950 series NPUs. The repository's README dates the initial release to September 30, 2026; its Python package identifies itself as version 0.1.0. It gives teams an open source starting point for matrix multiplication and related model operations on that hardware. The performance figures so far come from DeepSeek's own tests.

The release matters to engineers assessing an Ascend software stack because it exposes kernel source code and familiar DeepGEMM entry points, with a documented route to build the extension. Full-model performance and operating cost still require workload testing on the intended system.

## What is in the release

DeepGEMM is a library for the matrix operations used repeatedly in AI models. The [Ascend README](https://github.com/deepseek-ai/DeepGEMM-Ascend/blob/main/README.md) lists BF16, FP8 and FP4 GEMM kernels, grouped operations for mixture of experts models, MQA logits, a fused MegaMoE operation and an mHC prenorm kernel. Its [Python package exports](https://github.com/deepseek-ai/DeepGEMM-Ascend/blob/8491bbb4b8c02a094a2318965f50c70438a3e73c/deep_gemm/__init__.py) show dense and grouped GEMM functions, MQA functions, scaling factor layout helpers and controls for JIT compilation and AI core use. The repository is under the MIT license.

DeepSeek says the Ascend port is fully API compatible with its [upstream DeepGEMM](https://github.com/deepseek-ai/DeepGEMM#interfaces). The same `deep_gemm` package name and much of the public interface may ease a code migration, but the README identifies a consequential data-layout difference: Ascend packs each pair of UE8M0 scaling factors along K into an `int16` and stores the packed values in MN-major order. The package exports functions to transform scaling factors into that required layout. An application using quantized inputs still has to supply suitable data and validate its own results; the compatibility statement is the maintainers' claim, not a migration test by BIG CHANGE.

## What a developer needs to try it

The [documented requirements](https://github.com/deepseek-ai/DeepGEMM-Ascend#requirements) are an Ascend NPU, CANN 9.20 with the `bisheng` and `ld.lld` tools, `torch_npu`, Python 3.10 or newer, and a compiler and standard library supporting C++20 `<format>`. The maintainers say development and validation were on the Ascend 950 series. `tilelang` is declared as a package dependency for the HC prenorm kernel; `tree-sitter` and `tree-sitter-cpp` are needed to generate type stubs when building from source. The README does not pin a `torch_npu` version or establish validation on other Ascend devices.

The README instructs developers to clone the repository with its submodules, build the C++ extension through `develop.sh`, or install from the checked-out source with `pip install . --no-build-isolation`. Its [build configuration](https://github.com/deepseek-ai/DeepGEMM-Ascend/blob/8491bbb4b8c02a094a2318965f50c70438a3e73c/setup.py) imports `torch` and `torch_npu`, locates the Ascend toolkit through `ASCEND_HOME_PATH` or `ASCEND_TOOLKIT_HOME`, and links against Ascend and Torch NPU libraries. Evaluating the kernels therefore needs the appropriate hardware and toolchain. The repository includes [operation tests](https://github.com/deepseek-ai/DeepGEMM-Ascend/tree/8491bbb4b8c02a094a2318965f50c70438a3e73c/tests). BIG CHANGE inspected the documentation and code; it did not run the build or kernels.

## How far the benchmark goes

DeepSeek reports measurements on an Ascend 950DT with CANN 9.20, using `bench_msprof` with cold L2 cache. For one dense BF16 matrix multiplication shape, M=4096, N=7168 and K=16384, its [performance table](https://github.com/deepseek-ai/DeepGEMM-Ascend#performance) lists 2,229.9 microseconds, 431 TFLOPS and 99.8% of the hardware limit stated in that table. At the same shape, it reports 1,117.6 microseconds and 861 TFLOPS for FP8×FP8. These are kernel measurements at specified types and shapes, not measured model throughput. The table also gives grouped GEMM, MQA, MegaMoE and prenorm results under their own configurations; its MegaMoE values average eight ranks with expert parallelism of eight, top-k routing of six and one shared expert.

The [test code](https://github.com/deepseek-ai/DeepGEMM-Ascend/blob/8491bbb4b8c02a094a2318965f50c70438a3e73c/tests/common.py) checks GEMM outputs against generated references and uses an ACLNN cross-check for supported dense cases. It explicitly skips that supplementary ACLNN comparison for some shapes and recipes. Published code and tests make the claims inspectable, but the README's numbers have not been independently reproduced by BIG CHANGE. The release does not provide a like-for-like benchmark against another accelerator stack, a cost comparison, or an end-to-end service result.

## The big change

DeepSeek has released its DeepGEMM Ascend kernel code for the Ascend 950 series.

Developers can inspect and build familiar GEMM and related model operations for that hardware. The release documents CANN 9.20 and an Ascend-specific scaling-factor layout.

This gives Ascend 950 teams kernel code to evaluate against their own model shapes. DeepSeek reports near-limit results on one Ascend 950DT test setup; BIG CHANGE has not run the kernels.

Full-application performance, support for other Ascend devices and operating cost remain to be established.

## Sources & further reading

- [DeepGEMM Ascend repository and README](https://github.com/deepseek-ai/DeepGEMM-Ascend): The September 30 release note, prerequisites, interface claims and project-authored benchmark tables. The README is a primary source from the project's maintainers, not independent validation.
- [Package exports](https://github.com/deepseek-ai/DeepGEMM-Ascend/blob/8491bbb4b8c02a094a2318965f50c70438a3e73c/deep_gemm/__init__.py) and [build configuration](https://github.com/deepseek-ai/DeepGEMM-Ascend/blob/8491bbb4b8c02a094a2318965f50c70438a3e73c/setup.py): The version, public Python functions, dependency declaration and source-build linkage at the initial commit.
- [GEMM test harness](https://github.com/deepseek-ai/DeepGEMM-Ascend/blob/8491bbb4b8c02a094a2318965f50c70438a3e73c/tests/common.py): How the repository checks outputs and calls its profiler, including conditions where an ACLNN comparison is skipped. These files do not independently reproduce the published numbers.

## Sources

- [DeepSeek AI, DeepGEMM Ascend repository and README](https://github.com/deepseek-ai/DeepGEMM-Ascend) — Primary release record for interfaces, prerequisites, installation and project-authored benchmarks; the published figures are not independent results.
- [Initial public commit](https://github.com/deepseek-ai/DeepGEMM-Ascend/commit/8491bbb4b8c02a094a2318965f50c70438a3e73c) — Initial public commit dated 01:07 UTC, used to fix the source revision and verify release timing.
- [DeepGEMM Ascend Python package exports](https://github.com/deepseek-ai/DeepGEMM-Ascend/blob/8491bbb4b8c02a094a2318965f50c70438a3e73c/deep_gemm/__init__.py) — Version 0.1.0 and exposed GEMM, MQA, MegaMoE, scaling-layout and JIT controls; exports do not prove runtime operation on our systems.
- [Build configuration and GEMM test harness](https://github.com/deepseek-ai/DeepGEMM-Ascend/blob/8491bbb4b8c02a094a2318965f50c70438a3e73c/tests/common.py) — Test logic, correctness references and conditional ACLNN checks; code inspection is not a reproduced benchmark.
The BIG CHANGE newsletter

The big picture. At your pace.

Recent stories on AI and robotics, the shifts worth watching and practical ideas to use. Choose a daily briefing, weekly digest or monthly perspective.