Zhipu AI Bets on Domestic Silicon With 1GW Data Center, Acquisition to Break Free From Nvidia

Zhipu AI Bets on Domestic Silicon With 1GW Data Center, Acquisition to Break Free From Nvidia

Zhipu AI is executing the most aggressive compute-independence strategy seen among China's independent large-model companies, commissioning a 1-gigawatt domestic data center, acquiring compiler software firm Zhongke Jiahe for hundreds of millions of renminbi, and opening preliminary talks on custom AI chip development — a three-layer vertical integration push designed to permanently reduce its exposure to Nvidia Corp.'s supply chain.

The 1GW facility, first reported by Bloomberg, is already partially operational and runs exclusively on domestic AI chips, primarily supporting training and inference for Zhipu's GLM model series. The power footprint — sufficient to supply roughly 750,000 households — places it among the largest compute infrastructure assets held by any independent AI lab in China. Zhipu has not commented publicly on the data center's chip vendors, total card count, or capital expenditure. A person familiar with the matter confirmed the facility's existence to Bloomberg on condition of anonymity.

The disclosure arrives as Zhipu's financials expose a structural tension that self-owned infrastructure is intended to resolve: revenue nearly doubled while costs expanded even faster.


Surging Costs Force Zhipu to Move Up the Compute Stack

Zhipu's 2025 revenue rose 131.9% year-over-year to RMB 724 million (US$100.6 million), with cloud API and deployment revenue — the inference-heavy segment — surging 292.6% to RMB 190 million (US$26.4 million). Yet cost of sales climbed 213.3% over the same period, from RMB 137 million (US$19 million) to RMB 428 million (US$59.4 million), outpacing revenue growth by a material margin. The company attributed the cost acceleration explicitly to rising third-party compute service fees.

Research and development expenditure reached RMB 3.18 billion (US$441.7 million) in 2025, contributing to a net loss of RMB 4.718 billion (US$655 million), up from RMB 2.958 billion (US$410.8 million) in 2024. A significant portion of R&D spending was directed at external compute vendors and advanced training infrastructure — the precise cost centers that vertical integration aims to compress.

In Q1 2026, API call volume jumped approximately 400% quarter-on-quarter, prompting Zhipu to raise API pricing by 83%. The pricing move signals that demand-side leverage exists, but it does not eliminate the underlying unit-economics problem: every additional token generated consumes real silicon, power, and memory, regardless of what the customer pays.


Acquiring Zhongke Jiahe Targets the Software Gap Between Models and Chips

The hardware buildout alone solves only part of the problem. Domestic AI chips from vendors including Cambricon, Huawei's Ascend, and others do not natively run workloads optimized for Nvidia's CUDA ecosystem. Model migration requires rewriting operator libraries, memory management routines, and communication stacks — a process that can erode 30–50% of theoretical peak performance in practice.

Zhongke Jiahe, whose engineering team traces its lineage to the compiler laboratory at the Institute of Computing Technology under the Chinese Academy of Sciences, sits precisely at this software-hardware interface. Its capabilities span compilers, virtual instruction sets, runtime environments, and inference engines. The company's SigInfer inference engine claims, under controlled test conditions, to reduce inference latency by up to 74 times, increase throughput by up to 3 times, and improve energy efficiency by 1.46 times versus baseline configurations. Those figures derive from vendor benchmarks and cannot be assumed to replicate production-environment performance at Zhipu's scale — but even partial realization would materially reduce per-token compute costs.

Zhipu has completed the acquisition, according to people familiar with the deal. Neither party has issued a formal announcement. The transaction price has been reported by multiple Chinese media outlets as "hundreds of millions of renminbi," implying a figure in the RMB 200 million–RMB 900 million (US$27.8 million–US$125 million) range, though the precise sum has not been confirmed.

Critically, Zhongke Jiahe's team has prior compiler development experience across Loongson, Sunway, Cambricon, and Huawei Ascend — giving Zhipu a software abstraction layer capable of spanning multiple domestic chip architectures rather than locking into a single vendor's ecosystem.


Custom Chip Talks Signal a Longer-Term Architecture Bet

The third and most forward-looking element of Zhipu's strategy involves co-developing custom AI chips with domestic semiconductor design firms, a move first reported by The Information. The project remains at an early evaluation stage, with no partner, architecture, tape-out schedule, or volume production timeline disclosed.

The strategic logic, however, is structurally sound. Inference workloads — unlike pre-training — exhibit relatively stable compute patterns: fixed model architectures, predictable memory access profiles, and consistent precision requirements. These characteristics make inference a viable target for application-specific chip optimization, where custom silicon can outperform general-purpose accelerators on cost-per-token metrics by a factor of two to five times at scale, based on comparable deployments at U.S. hyperscalers.

Zhipu's implicit question — whether future chip architectures should be designed around GLM's specific inference requirements rather than the reverse — mirrors the logic that led Google to develop its Tensor Processing Units and Amazon to build Trainium and Inferentia. The difference is that Zhipu is pursuing this path under export-control constraints that prevent it from accessing Nvidia's H100 or H200 series, compressing its timeline for domestic alternatives.


Infrastructure Stress Tests Reveal Inference as the Operational Bottleneck

The urgency behind Zhipu's infrastructure push is traceable to a series of visible operational failures over the past six months. In January 2026, rapid user growth for GLM Coding Plan forced Zhipu to cap new daily purchase quotas at 20% of prior levels, with existing users experiencing concurrency throttling and degraded response speeds during weekday afternoon peak hours.

By March, following the launch of GLM-5 and its Coding Agent feature, daily API call volumes reached hundreds of millions. Coding Agent workloads impose disproportionate infrastructure pressure — they require reading extended code contexts, generating large token volumes continuously, and repeatedly invoking search, execution, and debugging tools. Inference system load runs an order of magnitude higher than standard chat completions.

The stress eventually manifested as apparent model quality degradation: users reported garbled output, repetition loops, and anomalous character generation from GLM-5 under high-concurrency conditions. After weeks of investigation, Zhipu traced the failures not to model weight corruption but to KV Cache mismatches and race conditions in inference state management under high load — a finding the company documented in a technical blog post titled Scaling Pain: Large-Scale Coding Agent Inference in Practice.

Following the GLM-5.2 launch, The Information reported that daily token usage on the Vercel platform grew 27 times in the first week. Comparable capacity crises hit Moonshot AI after the Kimi K3 launch — user requests exceeded cluster capacity within 48 hours, forcing a suspension of new C-tier subscriptions — and Alibaba's Qwen after the debut of Qwen3.8-Max, a 2.4-trillion-parameter model that triggered rate-limit errors across Token Plan, Qoder, and API entry points.


DeepSeek's Parallel Buildout Suggests an Industry-Wide Structural Shift

Zhipu is not moving in isolation. DeepSeek is simultaneously building out its own compute infrastructure at a data center in Ulanqab, Inner Mongolia, staffing a dedicated compute operations team, and adding chip design personnel to develop inference-optimized proprietary AI chips. The convergence of two of China's most technically credible independent AI labs on identical strategic priorities — self-owned hardware, domestic software stacks, custom silicon — constitutes a structural signal rather than a coincidence.

The model company's traditional role as a compute buyer is giving way to a new posture: compute operator, infrastructure modifier, and eventually, chip architecture co-designer. This transition carries significant capital implications. Self-owned data centers require upfront chip procurement, power and cooling infrastructure, network buildout, depreciation schedules, and long-term operations expenditure. If model demand disappoints or chip generations turn over faster than depreciation cycles, the fixed-cost base becomes a liability.

At sufficient scale, however, the calculus inverts. The value of owned infrastructure is not simply lower unit cost — it is supply certainty, software-hardware co-optimization, and the ability to scale without being subject to third-party allocation queues and spot-market pricing volatility.


Policy Tailwinds Accelerate Domestic Chip Adoption at Infrastructure Scale

The strategic and commercial rationale for Zhipu's buildout is reinforced by a policy environment that has shifted from subsidizing chip R&D to mandating domestic chip procurement at the infrastructure level. State-funded new data centers are now required to deploy domestic AI chips exclusively, according to media reports. Bloomberg has reported that a national AI compute infrastructure plan valued at approximately RMB 2 trillion (US$277.8 billion) over five years is under consideration.

China's 15th Five-Year Plan, covering 2026–2030, designates a national integrated compute network as a major state engineering project, with directives to deepen the East-Data-West-Computing initiative and build a multi-tier compute infrastructure system intended to make compute resources as accessible as electricity and water.

For domestic chip vendors — Cambricon, Huawei Ascend, Moore Threads, Biren Technology, and others — the policy shift represents a transition from fragmented pilot deployments to sustained, large-scale procurement cycles. For Zhipu and its peers, it means the domestic silicon ecosystem they are betting on will receive structural demand support independent of any single company's purchasing decisions.

The question for investors and supply-chain observers is no longer whether Chinese AI labs will build domestic compute infrastructure — that decision has been made. The question is how quickly the domestic chip ecosystem can close the performance gap with Nvidia's restricted products, and whether companies like Zhipu can convert infrastructure ownership into durable unit-economics advantages before the capital burden of self-funded buildout strains their balance sheets further.

Related Coverage:

Zhipu AI's ARR Hits $1 Billion, Surging 15-Fold in Six Months

Subscribe to ChinaBiz Insider

Don’t miss out on the latest issues. Sign up now to get access to the library of members-only issues.
[email protected]
Subscribe