Huawei Ascend Adapts DeepSeek V4 on Launch Day

Huawei Ascend Adapts DeepSeek V4 on Launch Day

Huawei Technologies announced full compatibility for DeepSeek V4-Pro and V4-Flash across its Ascend AI accelerator lineup within hours of the models' April 24 release, marking a strategic push to position domestic chip infrastructure as the go-to alternative for China's rapidly evolving large language model ecosystem.

The synchronized rollout underscores intensifying competition between Huawei's homegrown semiconductors and Nvidia's embargoed GPUs, as Chinese AI developers seek viable paths to scale trillion-parameter models. DeepSeek's latest iteration extends context windows from 128,000 tokens to 1 million — a 10x leap — while introducing sliding-window KV cache compression to slash memory overhead. Huawei's Ascend 950 and A3 supernodes now deliver inference latency as low as 10 milliseconds for the lighter V4-Flash variant, processing 4,700 tokens per second (TPS) on single cards under 8K input loads.

Performance Optimizations Target MoE Bottlenecks

Huawei's adaptation centers on three hardware-level upgrades tailored to mixture-of-experts (MoE) architectures. First, native support for FP8, MXFP8, and MXFP4 precision formats halves memory footprints while doubling compute throughput compared to conventional FP16 operations. Second, sparse memory access pathways specifically address MoE models' non-contiguous parameter loads, a chronic bottleneck in distributed inference. Third, shared memory pools between Vector and Cube processing units eliminate redundant on-chip data transfers, directly reducing end-to-end latency.

For the flagship V4-Pro model, Ascend 950 achieves time-per-output-token (TPOT) of approximately 20ms at 8K input scales, with decode throughput hitting 4,700 TPS per accelerator. The streamlined V4-Flash variant cuts TPOT to 10ms but maintains 1,600 TPS single-card throughput, reflecting architectural trade-offs optimized for real-time agent interactions and coding workflows.

A3 Supernodes Enable Cloud-Scale Deployments

Huawei's Atlas 900 A3 SuperPod and Atlas 800 A3 air-cooled clusters extend performance gains to multi-rack configurations. Built around flat architectures with unified memory addressing, the A3 series delivers point-to-point interconnect bandwidth of 784GB/s, supporting 32- to 384-card deployments suited for telecom carriers and financial institutions managing peak concurrent loads. In 64-card A3 clusters running DeepSeek V4-Flash via vLLM engines, Huawei logged 2,000+ TPS per card under 8K/1K input-output tests, with V4-Pro deployments entering production optimization phases.

The company's reliance on NAND-backed solid-state units (SSUs) for KV cache storage diverges from GPU-centric memory hierarchies, prioritizing cost per terabyte over raw DRAM bandwidth. This design choice reflects China's supply chain realities: while high-bandwidth memory (HBM) remains import-constrained, NAND flash production is domestically mature, enabling economical scaling to million-token contexts.

Developer Tooling Aims to Close CUDA Gap

To lower onboarding friction, Ascend CANN's new PyPTO programming paradigm replaces traditional GPU-style kernel coding with Python APIs that automate pipeline orchestration and memory management. Huawei claims DeepSeek V4-specific operators — including attention mechanisms, compression routines, and multi-head convolution — can now be developed in days rather than weeks. TileLang-Ascend, the accompanying open-source framework hosted on TileAI, provides dual-track interfaces for both performance engineers and application developers.

These tooling investments directly target Nvidia CUDA's entrenched ecosystem advantage. By generating performant kernels automatically and maintaining backward compatibility through a virtualized PTO instruction set, Huawei seeks to neutralize the switching costs that have historically locked AI teams into Nvidia stacks. Whether enterprise developers will embrace Ascend's Python-first approach remains an open question, particularly as DeepSeek's codebase still carries CUDA optimizations.

Market Implications for China's AI Stack

The day-zero compatibility announcement reflects deeper shifts in China's AI supply chain. Since October 2023 export controls severed access to H100 and A100 GPUs, domestic model builders have faced a binary choice: ration scarce legacy chips or retool around unproven alternatives. Huawei's ability to match DeepSeek V4's launch cadence — and deliver competitive inference metrics — suggests the Ascend roadmap is maturing faster than external observers expected.

For DeepSeek, Huawei's backing provides critical leverage in negotiations with cloud providers. While the startup's models remain framework-agnostic, Ascend-optimized deployments could unlock pricing advantages in China's intensely competitive AI-as-a-service market, where Alibaba Cloud, Tencent, and ByteDance's Volcengine vie for latency-sensitive workloads. The 1M-token context window, meanwhile, positions V4 for document-heavy enterprise use cases — legal review, financial analysis, codebase navigation — where retrieval-augmented generation has proven cumbersome.

Longer-term questions center on training scalability. Huawei provided reference implementations for Ascend A3-based training but disclosed no benchmark results. Given that DeepSeek trained its 671-billion-parameter V3 model on approximately 14.8 trillion tokens, validating Ascend's capabilities at that scale will determine whether China's AI infrastructure can sustain independent innovation cycles — or remain perpetually dependent on hoarded pre-sanction hardware.

Related Coverage:

DeepSeek Unveils V4 Preview With Million-Token Context Window

Subscribe to ChinaBiz Insider

Don’t miss out on the latest issues. Sign up now to get access to the library of members-only issues.
[email protected]
Subscribe