DeepSeek Ports DeepGEMM to Huawei Ascend 950 With BF16, FP8 and FP4 Kernels


DeepSeek released DeepGEMM Ascend on September 30, 2026, bringing its matrix-multiplication kernel library to Huawei's Ascend platform. The initial release targets Ascend 950 devices and supports BF16, FP8 and FP4 GEMM, along with MQA logits and MegaMoE operators. The repository is available under the MIT License.

DeepSeek says the Ascend implementation is API-compatible with DeepGEMM, allowing software built around the library's existing interfaces to retain the same development workflow. The port replaces the CUDA-oriented execution path with an implementation built around Ascend matrix multiply-add primitives and Ascend-specific data layouts, address calculations and pipelining.

The release gives developers a concrete open-source kernel layer for running DeepSeek-oriented workloads on Huawei's latest accelerator generation. It also expands the software available around Ascend beyond framework-level model support: GEMM and mixture-of-experts kernels sit directly on performance-critical execution paths for large-model training and inference.

What DeepGEMM Ascend includes

The first public release lists support for:

Capability DeepGEMM Ascend support
Target hardware Huawei Ascend 950
GEMM data types BF16, FP8, FP4
MQA logits Supported
MegaMoE Supported
API compatibility DeepGEMM-compatible
License MIT

The implementation uses Ascend-specific optimizations including sparse data loading and coroutine-based pipelining. DeepSeek describes the design goal as keeping kernel code compact while hiding lower-level details such as fractal layouts, alignment constraints and address calculations.

DeepSeek also publishes kernel measurements in the repository. These are project-reported microbenchmarks, not independent cross-platform results. They are useful for evaluating the maturity of the Ascend port, while comparisons with NVIDIA hardware require matched models, precision, shapes, software versions and power limits.

TileLang is part of the emerging Ascend kernel stack

The DeepGEMM Ascend repository credits TileLang for its mHC kernel and Huawei for technical support and engineering expertise. TileLang's separate Ascend implementation provides a Python-oriented DSL and compiler path for high-performance kernels on Ascend NPUs.

TileLang-Ascend currently documents two compiler routes, using Ascend C/PTO and AscendNPU IR. Its public documentation includes GEMM, vector and attention kernels and identifies Huawei Ascend A2 and A3 among tested accelerator families. The project has also published DeepSeek V4-oriented kernels during 2026.

That relationship matters because a usable accelerator ecosystem needs more than model checkpoints. Kernel libraries, compiler abstractions, communication libraries and framework integrations determine how much of the underlying hardware can be exposed to training and inference software without every team writing device-specific kernels from scratch.

Why the Ascend port matters

Reuters reported on September 30 that DeepSeek and Huawei are working on programming infrastructure for Ascend, including open-source compute and communication components. DeepGEMM Ascend is a directly inspectable artifact from that effort: the code is published under DeepSeek's GitHub organization and explicitly targets Ascend 950.

For teams evaluating Ascend hardware, API compatibility with the existing DeepGEMM interface can reduce the amount of application-level code that changes during hardware migration. Actual migration effort still depends on the rest of the stack, including framework support, collective communication, model-specific operators, CANN versions and deployment tooling.

The release is also relevant to the broader accelerator-software competition. NVIDIA's mature CUDA ecosystem combines compilers, optimized libraries, profiling tools, communication primitives and broad framework support. DeepGEMM Ascend addresses a narrower but important layer: performance-critical kernels used by modern large-model workloads.

Deployment considerations

The current release should be evaluated as an Ascend 950 kernel library, not as evidence that arbitrary CUDA applications can move unchanged to Ascend. DeepGEMM API compatibility covers this library's interface; other CUDA dependencies in an application retain their own porting requirements.

Teams considering the port should validate the exact model operators they need, supported numerical formats, CANN/toolchain requirements and kernel correctness on their target hardware. Project-reported benchmark numbers should be reproduced under the intended model shapes before they are used for capacity or cost planning.

For DeepSeek workloads specifically, the combination of DeepGEMM Ascend and the wider TileLang-Ascend work materially expands the publicly inspectable low-level software available for Huawei accelerators. The next useful measurements will be end-to-end training and inference results under reproducible configurations, where kernel performance can be separated from communication, memory and framework overhead.

Sources