DeepGEMM on Ascend hits 99.8% hardware utilization — rare for a new platform port
on: deepseek-ai/DeepGEMM-Ascend
DeepGEMM-Ascend ports DeepSeek's GEMM library to Huawei's Ascend NPU, and the headline claim is that it reaches near-peak hardware performance across a wide range of matrix shapes. The benchmarks on an Ascend 950DT back that up: BF16 dense GEMM hits 99.8% of the hardware limit, FP8 and FP4 mixed-precision variants land at 99.5% and 98.3% respectively. Those are not typical numbers for a new hardware port.
The engineering story behind those numbers involves two Ascend-specific techniques: sparse data loading and coroutine-based pipelining. The library wraps the Ascend MAD (matrix multiply-add) primitives behind a lightweight abstraction that hides fractal matrix layouts, alignment constraints, and address calculations — the kind of detail that normally forces kernel authors to write verbose, hardware-entangled code. The abstraction is thin enough that the kernels stay concise without sacrificing efficiency.
Beyond plain dense GEMM, the library covers the full set of operations DeepSeek's own models actually need: M-grouped GEMM for MoE expert routing, MQA logits for the Lightning Indexer, a fused MegaMoE kernel that combines EP dispatch, two grouped GEMMs, SwiGLU, and combine in a single pass, and an HC Prenorm GEMM for the mHC module. The MQA kernel is explicitly FIX-pipe bound rather than compute bound, saturating that pipe at 99% utilization — a useful signal about where the bottleneck actually lives on this hardware.
The MegaMoE benchmarks are particularly informative for anyone running large MoE inference. With 384 experts, hidden size 7168, and 16384 tokens, the fused kernel achieves over 846 TFLOPS of computation while keeping communication bandwidth around 103 GB/s per rank averaged across 8 ranks. That balance matters more than raw FLOPS when you are scaling expert parallelism.
The API compatibility story is straightforward: the package uses the same name as DeepGEMM, so existing code targeting NVIDIA hardware can target Ascend by swapping the installed package. The scaling factor format does differ — pairs of UE8M0 factors along the K dimension are packed into int16 and stored in MN-major order — so any code that touches scaling factors directly will need adjustment.
The dependency stack is specific: CANN 9.20, torch_npu, Python 3.10 or higher, and C++20 with <format> support. The mHC kernel also pulls in TileLang. Development and validation happened on the Ascend 950 series, so behavior on other Ascend generations is uncharted.
For teams already committed to Ascend hardware and running DeepSeek-style architectures, this is a credible path to hardware-limit performance without writing low-level NPU kernels from scratch.
A credible Ascend port of DeepGEMM that hits 99.8% hardware utilization on BF16 and covers the full MoE kernel stack DeepSeek models actually need.