Huawei Reengineers AI Infrastructure With Early Ascend 960 Launch
Huawei Technologies has overhauled its artificial intelligence computing architecture across silicon, server nodes, and physical data centers, signaling a strategic pivot from raw single-chip performance to system-level engineering amid the global race for 10-trillion-parameter AI models.
At the Huawei Connect 2026 conference in Shanghai, the Chinese technology conglomerate confirmed its next-generation Ascend 960 series will launch up to three quarters ahead of schedule in 2027. This acceleration provides critical capital expenditure (capex) certainty for domestic cloud providers and AI developers, while demonstrating Huawei’s reliance on its telecommunications heritage to bypass physical scaling limits.
Market observers note that as AI clusters expand toward 100,000-card scale, the industry bottleneck has decisively shifted. Huawei’s latest infrastructure strategy indicates that minimizing interconnect latency and optimizing data center thermal dynamics are now equally critical to gross floating-point operations per second (FLOPS).
Accelerating Silicon Shipments Shores Up Capex Certainty
The accelerated rollout of the Ascend 960 series serves as a stabilizing signal for a domestic AI supply chain requiring predictable product cycles to justify massive infrastructure investments. The training-focused 960DT will be ready by the first quarter of 2027, three quarters earlier than originally planned, while the inference-oriented 960PR will launch in the third quarter of 2027.
Hardware specifications released at the event highlight the system's focus on bandwidth. The 960DT delivers 2 PFLOPS of FP8 compute and 4 PFLOPS of FP4, supported by 288GB of on-chip HBM and memory bandwidth reaching 9.6 TB/s. Huawei has also codified a one-year iteration cycle, projecting the release of the Ascend 970 in 2028 and the 980 in 2029.
This forward guidance is underpinned by existing commercial traction. The company disclosed that deployment of its current-generation Ascend 910C super nodes has surpassed 1,000 sets, with the Ascend 950 also entering large-scale commercial use. By ensuring silicon availability, Huawei effectively sets the development timeline for downstream vendors spanning optical modules, liquid cooling, and network storage.
Overhauling Node Architecture Targets Interconnect Bottlenecks
To address the diminishing returns of massive computing clusters—where traditional architectures can lose more than 40% of training time to communication wait states—Huawei has restructured its server hierarchy. The company introduced a 4,096-card "Super Node" that functions logically as a single computer.
Simulation data from Huawei’s Markov Lab indicates that a 100,000-card cluster built on 4K super nodes yields a 2.75x increase in floating-point utilization compared to traditional 8-card server equivalents. This scaling is facilitated by transitioning from conventional pluggable optics to Near-Packaged Optics (NPO).
The newly unveiled Hi-ONE optical engine delivers a transmission capacity of 7.2T. In a standard Ascend 960 super node, 5,500 Hi-ONE units replace 48,000 traditional 800G optical modules. This integration slashes power consumption by more than 550 kilowatts and doubles the system's mean time between failures (MTBF), pushing availability to 99.8%.
To unify this hardware, Huawei deployed its UnifiedBus protocol, enabling unified memory addressing across physical servers. When paired with the new OceanStor M900 PB-level KV cache system for inference, this architecture supports theoretical multi-rail network topologies scaling up to one million cards.
Pioneering 3D Facilities Decouples Infrastructure Lifecycles
As rack power densities surge from legacy levels to upwards of 100 to 200 kilowatts, Huawei has entirely redesigned the physical data center. Following a pilot project in Wuhu, Anhui province, the company showcased a "3D Data Center" model that vertically stacks cooling systems at the base, IT equipment in the center, and power supply networks at the top.
This vertical stratification borrows operational logic from semiconductor fabrication plants. It physically isolates sub-systems, minimizing the risk of cascading failures while drastically shortening power and liquid-cooling routing. More crucially, it decouples the 18-to-24-month construction lifecycle of traditional data centers from the rapid one-year iteration cycle of AI chips.
By achieving a 90% prefabrication rate for mechanical and electrical systems, Huawei has compressed facility construction time from six months to three. The modular nature of the 3D architecture allows operators to upgrade IT equipment without overhauling existing power and cooling infrastructure, presenting a highly replicable, standardized product for capital-intensive AI hyperscalers.
Ultimately, Huawei's 2026 showcase minimizes geopolitical rhetoric in favor of engineering pragmatism. By synchronizing chip delivery schedules, optical engine bandwidth, and data center thermal dynamics, the company is attempting to define a proprietary, end-to-end standard for the next era of industrial AI computing.