China's AI Firms Are Rationing Tokens. That's the Bubble Warning Nobody Wants to Hear.

China's AI Firms Are Rationing Tokens. That's the Bubble Warning Nobody Wants to Hear.

China's artificial intelligence industry has hit a structural compute ceiling in 2026 — and the rationing of token capacity by the country's most advanced model companies signals a risk that goes well beyond supply-chain inconvenience: it raises the question of whether China's AI boom can survive its own cost structure.

The warning shot came not from a regulator or a rival, but from the product pages of China's best-funded AI labs. In June 2026, Zhipu AI, Moonshot AI's Kimi, and MiniMax — three companies that collectively represent the frontier of Chinese large language model development — all launched their most capable coding models to date, then almost simultaneously imposed purchasing restrictions on those same products. Zhipu's subscription plans now require daily queuing and have been repriced three times in a year. Kimi and MiniMax's APIs are running at persistent overload, with developers publicly waiting in line for token allocations. The irony is precise: the most capable AI products China has ever built are being sold on rationing terms more reminiscent of a planned economy than a technology platform.

The market implication is stark. When a software company limits how much of its product customers can buy, it has effectively reclassified its output from infinitely replicable code to an industrially constrained commodity — one with a hard production ceiling. That reclassification is the first structural alarm bell for China's AI investment cycle.


Yann LeCun's Cost Arithmetic Lands Hardest in Beijing

The theoretical framework for this crisis was articulated publicly on June 18, 2026, when Yann LeCun — widely recognized as one of the founding figures of modern deep learning — told CNBC that an AI bubble burst was not a distant scenario but an imminent one. His logic was disarmingly simple: the price of advanced AI products has been rising, but the cost of producing each token has not fallen fast enough to close the gap. Every major AI company is currently using investor capital to subsidize user consumption. If that cost curve doesn't bend sharply downward before the capital runs out, revenue will never reach the projections embedded in current valuations, and the collapse will follow.

The global data points supporting this thesis are not abstract. Elon Musk's xAI, which merged with SpaceX to reach a combined valuation of approximately $2 trillion, posted a quarterly loss of $2.5 billion against revenue of just over $800 million. Anthropic is spending $1.25 billion per month on compute — including renting Musk's own GPU clusters. OpenAI CEO Sam Altman has publicly acknowledged that costs are "a massive problem." These are not startups burning through seed rounds; they are the most capitalized AI entities in history, yet they remain structurally loss-making under current token economics.

For China, the same arithmetic applies — but with a harder constraint. U.S. companies operate within a domestic supply chain that includes Nvidia, AMD, Google TPUs, custom silicon from Microsoft and Meta, and virtually unlimited access to grid power and data center capacity. China's AI companies are operating within a supply chain that is simultaneously sanctioned, capacity-constrained, and architecturally immature.


Agentic Workloads Are Detonating Demand Faster Than Any Chatbot Ever Did

The proximate cause of the current crunch is not consumer chatbot traffic. The demand shock is being driven by AI coding agents and autonomous task-execution frameworks — the category that every major Chinese lab, and most global ones, identified in early 2026 as the primary commercial battleground.

The token economics of agentic workloads are categorically different from conversational AI. A single chat exchange consumes tens of thousands of tokens. A coding agent tasked with reproducing a research paper — as MiniMax demonstrated with its M3 model — ran autonomously for nearly 12 hours, consuming tokens at an estimated rate 50 to 100 times higher than a standard conversation. One U.S. developer calculated that his $200-per-month subscription to Claude and ChatGPT was consuming the equivalent of $2,048 in actual compute costs — a 10x subsidy ratio that is structurally unsustainable.

MiniMax responded to M3's launch by immediately abandoning its flat monthly subscription model in favor of per-token billing. Heavy users reported cost increases of 100% to 200% overnight. This is not a pricing anomaly — it is the industry's acknowledgment that software-style pricing logic no longer applies to AI inference. Token production is now priced like an industrial input, because that is what it has become.


Backers With Deep Pockets Are Discovering Their Own Pockets Have Holes

The conventional assumption was that China's AI startups were insulated from compute scarcity by their investor base. Zhipu AI counts Tencent, Alibaba, Ant Group, Meituan, and Xiaomi among its shareholders. Moonshot AI's largest shareholder is Alibaba, with approximately a 40% stake; Tencent is also a co-investor. On paper, having China's largest cloud computing operators as equity holders should translate into priority access to GPU capacity.

In practice, the scarcity is systemic enough to override those relationships. An ICT industry executive quoted in Chinese media described the hardware economics bluntly: two million renminbi (approximately $278,000) that previously purchased eight GPU servers now buys four or five, with vendors choosing to breach contracts rather than deliver at agreed prices. The shortage spans the entire stack — chips, high-bandwidth memory, advanced packaging, optical interconnects, and data center power capacity — and industry insiders estimate the tightness will persist for at least two more years.

In March 2026, Tencent Cloud raised prices on select Hunyuan model products by up to 400%, with Alibaba Cloud and Baidu Cloud following within hours. The message was unambiguous: even the cloud giants are rationing their own capacity.

More corrosively, at least one major cloud operator has publicly stated that scarce compute will be prioritized for its own highest-value workloads — meaning portfolio companies that depend on that operator for both funding and infrastructure now find themselves at the back of the queue. The investor, the landlord, and the competitor turn out to be the same entity. Zhipu AI's prospectus filings indicate that approximately 70% of its R&D expenditure goes toward compute procurement; the company has accumulated losses of approximately RMB 6.2 billion (US$861 million) over three and a half years. MiniMax carries a similar RMB 7 billion (approximately US$1.0 billion) loss over the same period, and its annual procurement commitment to Alibaba Cloud continues to rise.


DeepSeek Bets on Self-Sufficiency While Zhipu Optimizes Within Constraints

The compute crunch is forcing a strategic bifurcation among China's leading model companies. The common first move — adapting models to run on domestic hardware — has become near-universal. Zhipu AI trained its GLM-5.2 cluster on Huawei's Ascend processors, delivered exclusively through Shenzhou Digital using Ascend and Kuntai servers. Its GLM-Image multimodal model was the first top-tier Chinese multimodal system trained entirely on domestic silicon. DeepSeek went further: its V4-Pro release in April 2026 was delayed specifically to debut on Huawei Ascend, with its underlying inference code rewritten from Nvidia's CUDA framework to Huawei's CANN software stack — a signal that the company is actively decoupling from the U.S. chip ecosystem.

But the paths diverge sharply beyond that common ground. Zhipu AI continues to source its compute through cloud operators and shareholders. DeepSeek, by contrast, is moving directly upstream. The company is aggressively recruiting data center construction and operations talent, signaling plans to build gigawatt-scale proprietary compute infrastructure. In its June 2026 first funding round, founder Liang Wenfeng personally committed approximately RMB 20 billion (US$2.78 billion) as the largest individual contributor, structuring the round to exclude investors from board representation. The intent appears to be insulating an aggressive self-build strategy from shareholder interference — and, if successful, establishing DeepSeek as the first pure-play model company in China to own its own large-scale compute base.

On the efficiency side, Moonshot AI's Kimi is running inference on a proprietary architecture called Mooncake, which separates the prefill and decode phases of token generation and pools KV cache across entire GPU clusters for reuse — effectively extracting more requests from the same hardware. CEO Yang Zhilin has publicly framed "token efficiency" as the primary competitive variable, citing the company's MUON optimizer as evidence that training efficiency can be doubled without additional hardware. Zhipu AI's TileRT inference engine statically compiles entire computation graphs into persistent GPU kernels, pushing flagship model output to approximately 400 tokens per second; its ZCube network architecture, developed with Tsinghua University, reportedly improves inference throughput by 50%, reduces networking hardware costs by one-third, and cuts first-token latency by 40% — without adding a single GPU.

The price bifurcation in the market reflects this engineering divergence. DeepSeek cut its V4-Pro API price permanently to 25% of its original rate. Xiaomi's MiMo slashed prices by 90%. Tencent Cloud reduced its hosted DeepSeek cached-call pricing to RMB 0.025 per million tokens — cheaper than a domestic phone call. These cuts are concentrated in efficiency-optimized, cache-heavy workloads. Meanwhile, high-capability coding models from Zhipu and Kimi remain in short supply at rising prices. The market is not moving uniformly; it is stratifying by tier.


U.S. Competitors Operate From a Structurally Different Starting Position

The contrast with American AI infrastructure strategy is instructive, and it does not favor simple imitation. OpenAI's Stargate initiative commits $500 billion over four years to build 10 gigawatts of dedicated compute capacity, with Oracle, SoftBank, and other partners financing construction in exchange for long-term supply agreements. Microsoft continues to provide cloud infrastructure; CoreWeave and Oracle run parallel capacity; Broadcom is designing custom accelerator chips. OpenAI is not dependent on any single provider — it has constructed a captive infrastructure coalition in which every participant's revenue depends on OpenAI's continued scale.

Anthropic has taken a different but equally robust approach: a multi-cloud, multi-partner contract strategy. Its primary training infrastructure sits on Amazon Web Services, in a dedicated cluster of over one million chips whose capacity is reserved exclusively for Anthropic under a 10-year agreement valued at over $100 billion. It simultaneously holds a multi-billion-dollar TPU procurement commitment with Google, and supplements both with rented Nvidia capacity when needed. No single vendor has exclusivity; Anthropic retains control of model weights and pricing across all platforms.

Both strategies share one prerequisite: the most valuable components of the AI supply chain — advanced logic chips, high-bandwidth memory, mature hyperscale cloud platforms, custom accelerators, global data center capacity, and reliable grid power — are available domestically or through allied supply chains. U.S. AI companies are not managing scarcity; they are managing allocation within abundance. China's leading AI companies are managing scarcity within scarcity.


The Structural Question Remains Unanswered: Can Domestic Hardware Close the Gap?

Domestic chip market share data for 2025 shows Chinese AI processors reaching approximately 40% of domestic shipments, with Huawei holding nearly half of that share. Cambricon reported revenue growth of over 20 times year-over-year in 2025. These are not trivial numbers.

But the performance gap at the individual chip level remains significant. On multiple benchmark metrics, Nvidia's flagship GPUs outperform Huawei's Ascend by a factor of four to six. Huawei compensates at the cluster level by connecting hundreds of Ascend chips via optical interconnects into "super-nodes" that can match or exceed Nvidia cluster performance in aggregate — at the cost of approximately four times the power consumption per unit of compute output. The domestic hardware ecosystem currently offers a functional substitute, not a preferred alternative. And the entire domestic supply remains oversubscribed.

According to IDC projections, global annual token consumption in 2030 will exceed 2025 levels by more than 300 million times. Every optimistic analysis of China's AI market is built on demand-side projections of that magnitude. Almost none of them adequately model the supply-side question: which chips, in which data centers, powered by which grid connections, will produce those tokens — and at what cost per unit.

If that cost does not fall faster than the capital supporting China's AI ecosystem is consumed, the outcome Yann LeCun described is not a theoretical risk. It is a scheduled event. The current rationing of tokens by China's most capable AI companies is not a temporary inconvenience. It is the first audible signal from the bottom of the production pool.

Related Coverage:

Xiaomi Claims Viral “Hunter Alpha” Models as MiMo V2 Trio, Pressuring China’s Agent AI Pricing

DeepSeek Unveils V4 Preview With Million-Token Context Window

Subscribe to ChinaBiz Insider

Don’t miss out on the latest issues. Sign up now to get access to the library of members-only issues.
[email protected]
Subscribe