Huawei Ascend: How China Built Its Own AI Chip Ecosystem Under Sanctions

Huawei Ascend: How China Built Its Own AI Chip Ecosystem Under Sanctions

What Is the Huawei Ascend Program?

Huawei's Ascend is a family of AI processors — and the broader hardware-software ecosystem built around them — developed entirely in-house by the company's semiconductor division, HiSilicon. The program spans custom chip architecture (DaVinci), a proprietary operator library (CANN), a homegrown AI framework (MindSpore), and a cluster interconnect protocol (UnifiedBus/LingQu) designed to replace industry-standard technologies from Nvidia, Synopsys, and others.

What makes Ascend structurally significant is not any single chip, but the fact that it represents one of the few end-to-end AI compute stacks — from silicon to software — developed outside the United States. By 2026, the Ascend 950 had entered mass production, DeepSeek had listed Huawei Ascend NPUs alongside Nvidia GPUs in its technical reports, and Huawei's compute product line had surpassed its wireless division in R&D investment.


Why Did Huawei Build Its Own AI Chip?

The decision was driven by two converging forces: strategic foresight and existential necessity.

The foresight came first. In January 2016, HiSilicon president He Tingbo approached Nvidia CEO Jensen Huang at CES, asking to license mobile GPU technology. Huang declined. His reasoning — "I only know how to make GPUs, and I'll only focus on doing that well" — was a turning point. He Tingbo concluded that core technology "cannot be borrowed, cannot be begged for, and cannot be waited for."

Later that year, AlphaGo defeated Go champion Lee Sedol. Nvidia had already recognized that nearly all deep learning computation could be decomposed into matrix multiply-accumulate operations. In August 2016, Huawei's then-rotating chairman Xu Zhijun met with AI pioneer Yoshua Bengio in Ottawa; Bengio argued that deep learning would trigger the next industrial revolution. Xu returned convinced.

In June 2017, He Tingbo, HiSilicon Fellow Liao Heng, and Xu Zhijun held a pivotal meeting in Shanghai. The question: should Huawei build a dedicated AI chip? Opponents argued that application demand was unclear. Liao Heng countered that all future computing applications could eventually be expressed as neural network operations. Xu backed that view. The DaVinci project was formally launched.

The necessity came later — and harder. On May 16, 2019, the U.S. Commerce Department added Huawei to its Entity List, cutting off TSMC's 7nm EUV supply for Ascend 910A. On August 17, 2020, the Foreign Direct Product Rule (FDPR) was extended to cover any product incorporating U.S.-origin technology anywhere in its supply chain — effectively a 100% blockade. EDA tools, test instruments, international standards bodies: access closed across the board.

The sanctions removed any remaining internal debate. A company that had been hedging its bets was now forced to go all-in.


How Does the DaVinci Architecture Work — and Why Is It Different?

The core architectural decision was made by Liao Heng, who proposed a 3D Cube compute unit rather than the conventional 2D matrix approach. In a single clock cycle, a 16×16×16 three-dimensional structure completes an entire matrix operation — a design optimized not for minimum die area or peak power efficiency, but for maximum AI compute per unit of power.

This diverged sharply from the competing internal proposal, which aimed to replicate Nvidia's design philosophy as closely as possible. He Tingbo chose Liao Heng's approach.

Three other architectural decisions proved consequential:

Proprietary instruction set. Huawei rejected both Nvidia's CUDA path and ARM/RISC-V, building its own instruction set from scratch. The reasoning, articulated by chip solutions lead Jin Xi: at billion-user scale, any architecture built on a third-party instruction set will exhaust its optimization headroom precisely when it is needed most.

Broad-spectrum design. Liao Heng's "wide-spectrum" principle required a single architecture to scale across a price range spanning 10 million times — from sub-$5 edge chips to $500 million datacenter clusters. This forced generality into the design from day one.

Full-stack ownership. The compiler — described internally as the "holy grail" of the DaVinci project — required rebuilding from scratch after the first year's work was scrapped. The eventual compiler team numbered two to three hundred specialists. CANN (the operator library) was described by its architects as "the soul of Ascend — without CANN, the chip is just a piece of scrap metal."


What Did U.S. Sanctions Actually Force Huawei to Build?

The sanctions inadvertently created a comprehensive technology substitution program. Each restriction triggered a parallel development effort:

Blocked Technology

Huawei's Response

TSMC advanced node fabrication

Domestic wafer process development; yield improvement programs

EDA tools (Synopsys, Cadence)

Proprietary EDA tools — He Tingbo: “We built tools the U.S. doesn't even have”

Nvidia NVLink interconnect

LingQu (UnifiedBus) cluster interconnect protocol

PCIe/international bus standards

Forced after standards bodies revoked Huawei membership

High-end test instruments

Internal manufacturing of test equipment

Nvidia CUDA ecosystem

CANN + Ascend C (C++-based developer interface)

The LingQu interconnect origin is illustrative. It began after international standards organizations closed Huawei's membership access following the May 2019 sanctions. What started as a workaround became, in Liao Heng's description, the equivalent of "Qin Shi Huang unifying the six kingdoms" — not a patchwork of bus protocols, but a unified architecture that Huawei claims surpasses Nvidia NVLink in cluster-scale performance.

One unintended consequence: a chip originally sourced from the U.S. for connecting four Ascend 910 processors became unavailable. The team replaced it with a fully self-developed interconnect — which later became the foundation for Huawei's supernode systems.

Sanctions also had measurable blowback effects. Within roughly a year of implementation, Nvidia's revenue had declined by approximately $400 million. Nvidia's China business at the time represented roughly half of its total shipments.


Who Are the Key People, and What Did They Actually Do?

The Ascend story is inseparable from a small group of decision-makers who staked their careers — and in some cases their health — on the program.

He Tingbo, HiSilicon president since 2011, received the original "backup plan" mandate from Ren Zhengfei with a budget of $400 million and a team ceiling of 20,000 people (the actual core team stayed under 3,000). In 2021 — the hardest year — she described the situation as "a mountain lake being drained dry." By end-2022, she told those around her: "We have the steering wheel. We've basically survived." When asked whether Huawei could still function, she looked up from her meal and said simply: "Yes — and we've already built it."

Xu Zhijun, then rotating chairman, served as the project's internal political anchor. He chaired the DaVinci steering committee — whose members were uniformly heads of major business units — and repeatedly absorbed pressure from senior colleagues who questioned the program's enormous cost and uncertain returns. His decision framework: if a technology will reshape the industry and create customer value, concentrate resources and attack.

Liao Heng, HiSilicon's chief scientist and a Fellow of Huawei's 2012 Labs, provided the core architectural vision. His 3D Cube design, his broad-spectrum philosophy, and his LingQu interconnect work are the structural foundations of what Ascend became. His framing of the sanctions: "Necessity is the mother of invention. Being truly needed is the strongest driver of innovation."

Lei Jingchen, head of Huawei's Turing business unit, articulated the competitive strategy shift after sanctions: "We fight with systems, not individual chips." His hair turned completely white between 2018 and 2022. He declined multiple high-compensation offers from competitors.

Ren Zhengfei, Huawei's founder and CEO, applied his signature "Van Fleet ammunition" doctrine — concentrate superior resources to saturate a strategic breakthrough point — to the AI chip program. He named the AI combat team "Fourth Field Army," after the People's Liberation Army unit known for fighting hard battles and rapid advances. He has repeatedly cited Huawei's private ownership structure as what made decade-scale investment possible: "If we were a listed company, could we have survived?"


What Is the "System vs. Chip" Strategy?

After 2020, Huawei stopped competing on single-chip specifications — a contest it could not win given process node restrictions — and shifted to competing at the system level.

The strategic logic, articulated by Zhang Dixuan of the Ascend compute product team: "If we sell individual GPU cards, we directly expose our process node disadvantage. We compensate through 'area for compute, engineering for process' — closing the gap at the system level."

In practice, this means:

  • Density engineering: Huawei's 1,000-petaflop AI installation for Pengcheng Lab fits in 64 cabinets — equivalent to 500,000 desktop PCs.
  • Cluster efficiency: A 1,000-chip Ascend 910 cluster achieves approximately 90% linear scaling efficiency — roughly 15 percentage points above comparable market products at launch.
  • Full-stack control: Because Huawei owns the chip, interconnect, software stack, and system integration layer, it can optimize across all levels simultaneously. Zhou Mingyuan, an international GPU architecture expert who joined Huawei in January 2020, noted that Huawei's commercial product iteration cycle runs more than 50% faster than Nvidia's.
  • Supernode architecture: The Atlas 950 SuperNode, announced in July 2026, represents the current expression of this philosophy — compute delivered as a unified system rather than discrete accelerator cards.

The competitive framing internally: "Like a snake's long coiling body — where the opponent is strong, we avoid; where they are weak, we replace, locking on at every joint."


How Did DeepSeek Change the Equation?

The relationship between Huawei Ascend and DeepSeek represents a structural shift from reactive to proactive.

In summer 2024, HiSilicon engineers met DeepSeek founder Liang Wenfeng for the first time in Shanghai — he arrived in a backpack and sneakers, sitting quietly at the edge of the conference table. By November 2024, Liang personally called to propose deploying DeepSeek-R1 on Ascend A3 SuperPoD, offering to open-source code and co-develop. On the night of December 3, 2024, in a Guangzhou hotel, Liang set a demanding target: 100 PFLOPS on the Ascend A3 SuperPoD. Huawei's engineers accepted.

Over the 2025 Lunar New Year, 230-plus Huawei engineers organized into five task forces — architecture, resource provisioning, Ascend operators, model optimization, and operations — deployed over 3,000 Ascend chips at Huawei Cloud's Guian data center, bringing DeepSeek-V3/R1 online within three days.

The outcome: DeepSeek's V4 technical report listed two compute platforms — Nvidia GPU and Huawei Ascend NPU. Ren Zhengfei described the DeepSeek moment as the first time he genuinely felt that Chinese companies had a real chance at true software-hardware co-optimization.

The structural significance: DeepSeek's decision to adopt Ascend as a primary compute platform was not driven by regulation or subsidy, but by performance validation. It provided external proof-of-concept for the system-level strategy.


What Are the Key Constraints and Variables Going Forward?

Process node gap: Ascend chips are manufactured at domestic Chinese fabs at process nodes behind TSMC's leading edge. Huawei's system-level engineering compensates partially, but the gap remains a structural constraint on raw single-chip performance.

Software ecosystem maturity: Nvidia's CUDA ecosystem has a roughly 15-year head start and deep integration with PyTorch and TensorFlow — the "Wintel" of AI compute. Huawei's CANN opened fully in August 2025. Ascend C (the C++-based developer interface) lowered the barrier to entry significantly, but ecosystem depth takes time to accumulate.

Supply chain dependencies: Despite substantial progress, domestic Chinese semiconductor supply chains for advanced packaging, photolithography, and materials remain in development. He Tingbo spent nearly half her time in 2021–2022 coordinating domestic supply chain partners — some of whom proved unreliable under sanctions pressure.

Scale of investment: Huawei's R&D spending rose from ¥131.7 billion (15.3% of revenue) in 2018 to ¥161.5 billion (25.1% of revenue) in 2022 — increasing in absolute terms even as revenue declined. This level of sustained investment is structurally enabled by Huawei's private ownership, which insulates it from quarterly earnings pressure.

Revenue trajectory: A Reuters report in March 2026 cited customer test completions and orders for the Ascend 950 PR chip, with estimated 2026 revenue of ¥37.5–52.5 billion. The compute product team now exceeds 15,000 people; 2026 R&D investment surpassed the wireless product line for the first time.


What Comes Next?

Several structural trajectories are visible:

Iteration normalization: Huawei announced the Ascend 950 in September 2025 and previewed the 960 and 970 at the same event — signaling a return to annual chip generation cadence after years of disruption.

Ecosystem expansion: CANN's full open-sourcing in August 2025 is a direct bid to accelerate third-party developer adoption. The "Tao Law" — a methodology synthesized from 381 chip designs across 2020–2026 — provides a systematic framework for future development cycles.

The two-ecosystem question: The most consequential long-term variable is whether global AI compute bifurcates into two largely separate ecosystems — Nvidia CUDA and Huawei Ascend — or whether one achieves sufficient cross-market penetration to become the dominant standard. DeepSeek's dual-platform support suggests the bifurcation scenario is already underway.

China domestic market dynamics: With U.S. export controls tightening in October 2023 and beyond, Chinese cloud providers, AI labs, and enterprise customers face structural pressure to qualify Ascend as a primary platform. The Ascend 910B reached 80% of Nvidia A100 performance in certain workloads for iFLYTEK — a benchmark that crossed an important threshold for enterprise adoption.

What began in 2017 as a hedge against dependency has become, by 2026, the primary compute infrastructure for a significant portion of China's AI industry. The "backup plan" is now the main plan.

Related Coverage:

Huawei’s Tao’s Law V2 Bypasses EUV Constraints, Repricing China’s Chip Supply Chain

Subscribe to ChinaBiz Insider

Don’t miss out on the latest issues. Sign up now to get access to the library of members-only issues.
[email protected]
Subscribe