Zhipu Undercuts DeepSeek With GLM-5.3 Flash on 100,000+ Domestic Chips

Zhipu Undercuts DeepSeek With GLM-5.3 Flash on 100,000+ Domestic Chips

China's Zhipu AI fired a direct shot at DeepSeek's pricing dominance on Wednesday, launching GLM-5.3 Flash — a 300-billion-parameter lightweight flagship model running entirely on a cluster of more than 100,000 domestic chips — at rates that undercut its rival across nearly all standard inference workloads.

The launch crystallizes a strategic inflection point in China's AI model market: as DeepSeek raised prices sharply in August 2026, citing compute scarcity, it inadvertently vacated a high-value demand segment that Zhipu is now moving aggressively to capture. The timing is deliberate. According to LatePost, DeepSeek V4 Flash's call volume on third-party developer platform OpenCode fell by half following its price hike, creating a measurable demand gap that Zhipu's new model is priced to absorb.

Market reception ahead of the official launch was unambiguous. Operating anonymously under the codename "Ox Alpha" on OpenRouter and OpenCode for five days prior to launch, the model accumulated more than 50 trillion tokens (50T) of traffic — entirely free of charge — shattering traffic growth records on both platforms.


Domestic Chip Cluster Draws SemiAnalysis, Pressures Nvidia's Moat

The hardware story behind GLM-5.3 Flash may carry more long-term significance than the model itself. Zhipu confirmed in its technical documentation that inference for the model is handled by a cluster exceeding 100,000 domestic AI chips, with the company asserting that "hardware efficiency and per-token cost have reached parity with mainstream Nvidia GPUs."

Semiconductor research firm SemiAnalysis responded swiftly on X, writing: "All traffic runs on domestic chips, with hardware efficiency and per-token cost now comparable to Nvidia GPUs. Following yesterday's announcement of Jalapeño [OpenAI's in-house inference chip], the CUDA moat faces another test."

According to LatePost, the chip suppliers likely include Huawei, Moore Threads, and Hygon, though Zhipu declined to comment on supplier identities and its technical documentation omits specific chip models. What the documentation does detail is the engineering stack required to make domestic silicon competitive at scale: a custom inference engine built on SGLang, W8A8 quantization with INT8/FP8/BF16 mixed-cache quantization, intra-node tensor parallelism, and a production-grade Encode–Prefill–Decode (EPD) disaggregated architecture. Zhipu claims this stack delivers a 3x end-to-end throughput improvement over the baseline on identical hardware.

The 100,000-chip deployment threshold is significant. SemiAnalysis noted that sustaining 100 trillion tokens per day of free inference — the volume recorded during the anonymous testing phase — was previously assumed to require only the most well-resourced frontier labs. That Zhipu achieved this on domestic hardware challenges a persistent assumption about the operational ceiling of non-Nvidia infrastructure.


Architecture Departs From Distillation Orthodoxy, Adds Native Multimodal

GLM-5.3 Flash carries 300 billion total parameters with 18 billion activated per forward pass, representing roughly 40% of the parameter footprint of its full-size sibling GLM-5.3 — a ratio that mirrors DeepSeek V4 Flash's positioning relative to its own flagship. Compared with GLM-4.5 (355B total, 32B activated, 92 layers), the new model compresses to 45 layers and 18B activated parameters, nearly halving both depth and activation cost.

Critically, Zhipu chose independent architecture design over the industry-standard distillation approach — where smaller models are trained by compressing larger ones. GLM-5.3 Flash employs what the company describes as the first open-source frontier model using a hybrid sparse-attention and linear-attention architecture. Versus GLM-5.3, attention computation volume drops 3.01x and KV cache size shrinks 4.44x, directly translating into lower per-token inference cost.

The model also marks Zhipu's return to multimodal capability — supporting image and video input — the first such feature since the company pivoted its strategic focus toward coding applications. This is not a marginal addition: native multimodality, rather than bolt-on support, broadens the addressable workload base and positions GLM-5.3 Flash to compete for enterprise use cases that pure-text models cannot serve.

On the Artificial Analysis Intelligence Index, Zhipu's internal benchmarks place GLM-5.3 Flash at 57 points — above the previous flagship GLM-5.2, level with Anthropic's Claude Opus 4.8, and ahead of DeepSeek V4 Pro's official score of 53.


Pricing Arithmetic Exposes DeepSeek's Post-Hike Vulnerability

The commercial logic of GLM-5.3 Flash is straightforward and aggressive. At RMB 0.8 per million input tokens and RMB 2.8 per million output tokens (approximately US$0.11 and US$0.39 respectively at the prevailing RMB 7.2 exchange rate), the model prices at exactly one-tenth of GLM-5.3. A two-week introductory half-price promotion cuts that further to one-twentieth of the full flagship.

The comparison against DeepSeek V4 Flash post-hike is stark. DeepSeek now charges RMB 1.5 per million input tokens and RMB 4.5 per million output tokens at off-peak rates, rising to RMB 9 per million output tokens during peak hours. GLM-5.3 Flash undercuts those figures in virtually all standard call patterns.

The gap against Western frontier models is wider still. Claude Opus 4.8 carries an official output price of US$25 per million tokens; GLM-5.3 Flash's international output price is US$0.50 — a 50x differential.

DeepSeek's August 2026 price increases — V4 Pro output rising from RMB 6 to RMB 13.5 (peak: RMB 27), V4 Flash output from RMB 2 to RMB 4.5 (peak: RMB 9) — were justified by the company on grounds of compute scarcity and cash flow preservation. A leaked audio recording attributed to DeepSeek founder Liang Wenfeng had previously suggested AI demand was inelastic. The OpenCode volume data suggests otherwise: demand for lightweight flagship-tier models is highly price-sensitive, and the market penalized DeepSeek's hike immediately.


China's Model Market Reprices Toward Commodity Infrastructure Logic

The broader pattern since June 2026 reveals a structural dynamic. Zhipu released GLM-5.2 on June 13; Moonshot AI launched Kimi K3 in mid-July; DeepSeek V4 Pro's official release followed in August; and Zhipu pushed GLM-5.3 on August 14 before now releasing the Flash variant. Each release compressed the performance-per-dollar ratio further.

Kimi K3's output price of RMB 100 per million tokens — more than triple its predecessor K2.6 — and Zhipu's own first-half 2026 price increases to RMB 28 per million output tokens now look like a brief window of pricing power that the market is already closing. China's AI model sector faces structural constraints that prevent frontier-style scarcity pricing: an abundance of competing models, active open-source deployments, and cloud vendor token subsidies collectively prevent any single model from commanding sustained premium positioning.

The economics of model inference, stripped of capex, are attractive — high gross margins on a per-token basis, provided volume and capability thresholds are met. GLM-5.3 Flash's pre-launch anonymous testing, which generated 50T tokens of demand in five days at zero price, confirms that the volume threshold is reachable. Whether Zhipu can convert that traffic into durable paid call volume — and whether its domestic chip stack can scale without the reliability and toolchain depth of Nvidia's ecosystem — remains the operative question for investors watching China's AI infrastructure buildout.

Related Coverage:

Zhipu AI Hits 7 Million API Users, Deploys 50,000 Domestic Chips as ARR Surges 15-Fold

Subscribe to ChinaBiz Insider

Don’t miss out on the latest issues. Sign up now to get access to the library of members-only issues.
[email protected]
Subscribe