Huawei Rewrites AI Interconnect Rules With Peerium Architecture and UnifiedBus, Claiming China Market Lead Over Nvidia
Huawei Technologies has unveiled a system-level computing architecture it believes renders single-chip benchmarking obsolete, staking a claim to global leadership in large-scale AI interconnect at a moment when the race to train 40-trillion-parameter models is redefining what infrastructure actually means.
At Huawei Connect 2026 in September, Rotating Chairman Xu Zhijun and HiSilicon Chief Scientist Liao Hengdisclosed the company's two foundational bets: the Peerium computing architecture and the UnifiedBus interconnect protocol. The announcements landed against a backdrop of persistent U.S. export controls that have blocked Huawei's access to leading-edge fabrication nodes — constraints the company is now openly reframing as a design forcing function rather than a ceiling. In the most pointed market-share claim of the event, Xu stated that based on measurable data, Huawei's Ascend series has already surpassed Nvidia in China's AI accelerator market.
The timing is not incidental. With China's frontier AI laboratories preparing to train models in the 10-trillion-to-40-trillion parameter range over the next two years — a compute demand that Liao estimates will require clusters of 200,000 to 256,000 cards — the addressable market for whoever can deliver reliable, high-MFU (Model FLOPs Utilization) supernodes at scale is enormous. Xu was blunt about the binding constraint: "The key going forward is not how hard we push, but how much we can supply."
Peerium and UnifiedBus Dismantle the Multi-Protocol Bottleneck
The central engineering problem Huawei is attacking is well understood in the industry but poorly solved. As GPU clusters scale from hundreds to hundreds of thousands of cards, communication latency and protocol-switching overhead cause aggregate compute efficiency to collapse — what engineers describe as a cliff-edge drop in effective throughput. Nvidia's answer has been architecturally bifurcated: NVLink for intra-rack scale-up bandwidth (reaching 1.8 TB/s in the NVL72 configuration) and InfiniBand for inter-rack scale-out connectivity. The penalty is stark — bandwidth falls from 1.8 TB/s inside the rack to roughly 0.2 TB/s the moment data crosses a cabinet boundary. Protocol conversion between the two networks introduces microsecond-level latency at every handoff.
UB eliminates that handoff entirely. By treating the entire network — from scale-up to scale-out — as a single unified bus protocol analogous to a computer's internal memory bus, Huawei claims data packets traverse the scale-up-to-scale-out boundary in approximately 150 nanoseconds, an order-of-magnitude improvement over conventional protocol-conversion overhead. Inter-rack bandwidth under UB reaches up to 800 GB/s. The practical result, as demonstrated in the 910C supernodes already delivered — Xu cited more than 1,000 units of the 384-card configuration shipped — is that physically distributed chips across multiple cabinets behave logically as a single computer.
Liao offered a precise analogy: "If UB is the crystallization of a group's wisdom, it was conceived by computer architects but raised by network technology." The protocol's development was deliberately housed within Huawei's computing division rather than its traditional networking unit, a structural choice Xu said was essential to achieving the low-latency, high-speed characteristics that telecom-grade IP protocols cannot deliver.
The Peerium architecture underpins UB at the software-hardware interface. Through nested parallelism (what Huawei terms "Nested BSP") and unified memory addressing, Peerium breaks from the von Neumann master-slave model, enabling strong scaling across millions of processors. Liao published the architecture in academic paper form — a deliberate signal, Xu suggested, that Huawei views Peerium as the canonical computing architecture for the AI era rather than a proprietary moat. The UB protocol has similarly been opened to the industry; Xu noted that the largest volume of post-release downloads came from Silicon Valley.
Hi-ONE Module Bets on NPO Over CPO, Targeting 40% Cost Reduction
Alongside the interconnect announcements, Huawei introduced the Hi-ONE optical module based on NPO technology. The module supports 7.2 Tb/s bandwidth with the optical engine positioned approximately 5 centimeters from the host chip — the only remaining copper segment — making the 950-series supernode effectively all-optical.
The NPO-versus-CPO (Co-Packaged Optics) debate has been a live fault line across the industry. Huawei's position is unambiguous: NPO delivers cost savings of approximately 40% relative to CPO implementations, while preserving a critical serviceability advantage — a failed optical engine does not require scrapping the entire AI accelerator. Xu characterized the choice as "an engineering pragmatism and cost trade-off," noting that internal consensus formed early given Huawei's integrated position across optical components, optical modules, and Ascend chips.
Industry validation has followed. At the most recent OIF (Optical Internetworking Forum) standardization discussions, opposition to NPO has dwindled to a handful of holdouts, with 40 supply-chain companies now backing the specification. Liao was candid about the manufacturing difficulty: thermal sensitivity in optical components means minute temperature variations alter diode characteristics, refractive indices, and waveguide performance. "The path to mass production was extremely arduous," he said. Xu projected that no other vendor would be able to replicate Hi-ONE in volume within the next two to three years.
CUDA Moat Eroding as Frontier Labs Shift Programming Paradigms
Perhaps the most strategically significant disclosure came from Liao's assessment of the software competitive landscape. For years, the CUDA ecosystem — the accumulated weight of academic research, PyTorch integrations, and developer muscle memory built around Nvidia's programming model — has been the decisive non-hardware barrier to Ascend adoption. Liao argued that this moat is actively dissolving, driven by a shift in how frontier models are built rather than by any direct Ascend-side improvement.
Over the past 18 months, he said, model architectures have become sufficiently fine-grained and mathematically complex that pure PyTorch-on-CUDA workflows can no longer maintain efficiency. Frontier labs are migrating toward large-scale kernel programming styles that require new compiler languages — languages that abstract away the underlying hardware and, in doing so, level the playing field between Ascend and Nvidia silicon. "I believe we are approaching or have reached parity," Liao said. "The gap has narrowed dramatically. When developers use these processors today, they no longer feel unfamiliar."
This dynamic creates what amounts to a structural window for domestic Chinese chip adoption: the switching cost of moving from CUDA is declining precisely as model complexity makes staying on CUDA increasingly expensive in engineering terms.
950DT Supernodes Target China's Six-to-Seven Frontier Lab Cohort
Xu provided the clearest public timeline yet for the Ascend 950DT training ramp. The 950 series launched first in a PR (production release) configuration optimized for inference, with limited initial volumes. The 950DT — the training-focused variant built on UB-enabled supernodes — remains in customer testing as of the conference, with volume supply expected by year-end 2025 or early 2026. "Starting next year, I believe a large volume of model training will be built on 950DT supernodes," Xu said.
The demand context Liao outlined is specific: China will have approximately six to seven frontier AI laboratories within two years, each targeting foundation language models in the 10-trillion-to-40-trillion parameter range. The 200,000-card supernode scale is not arbitrary — it is calibrated to the memory capacity required for those parameter counts and constrained by the power delivery architecture of Chinese national-grid-connected data center campuses, where Xu described 200,000 cards as "a relatively conservative number" for a single regional site.
On overseas expansion, Xu was direct: domestic demand is so far in excess of current supply that a broad international push is not on the table. A small number of countries have received test allocations and limited supply, but the governing principle is domestic-first. The Malaysia sovereign AI deployment reports — involving 910C-based infrastructure — Xu framed as geopolitically motivated multi-vendor diversification by customers, not a deliberate Huawei export strategy.
Supply Chain, Not Technology, Now Defines the Constraint
In a signal that shifts the competitive narrative from capability to execution, Xu identified supply chain capacity — not chip design — as the primary variable determining Huawei's market trajectory. He drew an explicit parallel to global semiconductor supply dynamics: the premium pricing currently enjoyed by optical module manufacturers and memory suppliers reflects demand that the entire industry failed to anticipate, not product differentiation. He projected that global supply-demand balance in AI infrastructure will not be achieved until approximately 2029, with China reaching equilibrium somewhat later given additional domestic supply-chain development requirements.
For investors tracking the Ascend ecosystem, the implication is a sustained period of supply-constrained revenue rather than demand-constrained growth — a structurally different risk profile than typical semiconductor cycles. Xu's call for the domestic supply chain to "accelerate capacity expansion, move faster, produce more" reads as a direct signal to Chinese semiconductor equipment, advanced packaging, and optical component suppliers.
The 960-series chip, which Xu confirmed is on a faster release cadence with specifications consistent with prior public disclosures and incorporating HBM (High Bandwidth Memory), represents the next hardware node in the Peerium roadmap. "Without progress in HBM, there is no 960," Xu said — a rare public acknowledgment of the HBM supply dependency that sits beneath Huawei's training chip ambitions.
Related Coverage:
Huawei's Aito Zhijie RX Bets on L3 Architecture as China Legalizes Conditional Autonomy