Huawei's Super-Node Architecture: A Technical Deep Dive Into China's AI Infrastructure Ambitions
At Huawei's Connect 2025 conference, the Chinese tech giant unveiled what may be its most ambitious AI infrastructure play yet: the Atlas 950 and Atlas 960 SuperPoD super-nodes, supporting 8,192 and 15,488 Ascend cards respectively. According to a recent analysis by Huaxi Securities, these developments represent more than incremental improvements—they signal a fundamental reimagining of AI cluster architecture that could challenge Western dominance in high-performance computing.
The Scale Advantage
The numbers tell a compelling story. Compared to NVIDIA's forthcoming NVL144, slated for release in the second half of next year, Huawei's Atlas 950 super-node delivers a scale advantage of 56.8x, total computing power of 6.7x, and memory capacity 15 times larger at 1,152TB. Perhaps most striking is the interconnect bandwidth advantage: 16.3PB/s versus NVIDIA's offering—a 62x improvement.
Yet raw specifications only tell part of the story. The real innovation lies in Huawei's solution to what the company identifies as two critical technical challenges in super-node interconnection.
Solving the Distance-Reliability Paradox
Current electrical interconnect technologies suffer from distance limitations at high speeds, supporting at most two-cabinet interconnections. Optical interconnect technologies can bridge longer distances across multiple cabinets but fail to meet reliability requirements. Huawei's approach introduces what it calls "hundred-nanosecond-level" fault detection and protection switching mechanisms, combined with redesigned optical components and interconnect chips.
The bandwidth-latency challenge proves equally formidable. Current cross-cabinet card interconnect bandwidth falls short of super-node requirements by a factor of five, while cross-cabinet latency currently maxes out around 3 microseconds—24% above Atlas 950/960 design targets. When latency operates in the 2-3 microsecond range, approaching physical limits, even 0.1 microsecond improvements represent significant engineering challenges.
The "LingQu" Protocol Innovation
Huawei's solution centers on a new interconnect protocol dubbed "LingQu" (UB, UnifiedBus). Through multi-port aggregation and high-density packaging technology, combined with peer-to-peer architecture and unified protocols, the system achieves TB-level ultra-high bandwidth and 2.1 microsecond ultra-low latency.
The LingQu 1.0-based Atlas 900 super-node (CloudMatrix 384) began delivery in March 2025, with over 300 commercial deployments to date. The enhanced LingQu 2.0 protocol not only optimizes performance and scales up capabilities but will be opened to industry partners—a move that could accelerate ecosystem adoption.
Massive Scale Ambitions
Based on LingQu 2.0, Huawei simultaneously launched the Atlas 950 SuperCluster, a 500,000-card cluster comprising 64 interconnected Atlas 950 super-nodes. This system integrates over 520,000 Ascend 950DT cards, delivering 524 EFLOPS in FP8 total computing power.
The networking architecture employs UB-Mesh technology with nD-FullMesh topology, prioritizing short-range direct interconnect paths to minimize data movement distances and reduce switch usage. Within racks, 2D-FullMesh networking is used, while inter-rack connections utilize single-layer UB Switch interconnects, enabling linear scaling from 64 cards to 8,192 cards.
Looking Ahead: Million-Card Clusters
By Q4 2027, Huawei plans to launch the Atlas 960 SuperCluster based on Atlas 960 super-nodes, scaling to million-card cluster levels with 2 ZFLOPS FP8 computing power and 4 ZFLOPS FP4 computing power.
The company has also mapped out its Ascend 950-970 series chip roadmap through 2028, with the Ascend 950 series targeting inference and training scenarios, the Ascend 960 series doubling key specifications, and the Ascend 970 series planned for Q4 2028.
As geopolitical tensions continue reshaping the semiconductor landscape, Huawei's super-node architecture represents more than technological advancement—it's a strategic pivot toward indigenous AI infrastructure capabilities that could redefine competitive dynamics in high-performance computing.