ChinaBiz Briefing | Alibaba's Qwen3.8-Flash, Zhipu’s Chip Bet, Li Auto’s Margin Crunch

ChinaBiz Briefing | Alibaba's Qwen3.8-Flash, Zhipu’s Chip Bet, Li Auto’s Margin Crunch

China's AI sector is entering a new phase — one defined less by benchmark races and more by the harder questions of unit economics, infrastructure capital, and pricing power. Wednesday's news flow delivers a rare convergence: a DeepSeek financial disclosure that reframes the company as a capital markets story, an Alibaba model launch that structurally compresses inference costs, a Zhipu pricing offensive enabled by domestic chips, a MiniMax earnings beat that masks deepening margin stress, and a Li Auto quarter that forces investors to weigh a credible technology pivot against a profitability trough. Together, they sketch the contours of an industry moving from land-grab to monetization — unevenly, and under pressure.


DeepSeek Posts 82.9% API Margin, Eyes $69.4B Valuation and 2027 IPO

DeepSeek generated approximately RMB 475 million (US$65.9 million) in revenue across the first seven months of 2026 — roughly ten times its full-year 2025 total — while posting an API gross margin of 82.9%, according to figures cited by The Information and Bloomberg on Aug. 26. Net losses narrowed to RMB 715 million for the same period, down from RMB 935 million for all of 2025, even as infrastructure spending surged nearly tenfold to RMB 11 billion. A second funding round targeting RMB 50 billion at a pre-money valuation of RMB 500 billion (US$69.4 billion) is expected to close before month-end; investment banks are already on retainer for a STAR Market IPO filing targeted by end-2026, with a listing expected in 2027.

The 82.9% API margin — against OpenAI's 39% and Anthropic's projected 63% for full-year 2026 — is the figure that will anchor institutional debate, though DeepSeek's blended margin of 44.6% is the more relevant near-term profitability indicator. The company's August price hikes, which lifted peak-hour output rates to US$3.96 per million tokens from US$0.87, signal that management believes developer lock-in is now sufficient to absorb cost increases — a hypothesis Zhipu's launch (see below) is already testing. For China's domestic chip supply chain, DeepSeek's infrastructure ramp from RMB 1.2 billion in 2025 to RMB 11 billion in seven months represents a meaningful demand signal as U.S. export controls continue to constrain Nvidia H100 and B200 access.


Alibaba's Qwen3.8-Flash Prices at 3% of Claude Opus, Doubles as Qwen4 Blueprint

Alibaba on Aug. 26 released Qwen3.8-Flash, a 125-billion-parameter mixture-of-experts model that activates only 6 billion parameters per token, priced at RMB 1 per million input tokens and RMB 3 per million output tokens (approximately US$0.14 and US$0.42) — 3% of Claude Opus 4.6's equivalent rate. On SWE-bench Pro, the model scores 58.7 against 16.5 for the prior-generation Qwen3.7-Plus, which carries three times the activated parameters. On CoWorkBench and JobBench, it outperforms Claude Opus 4.6 by 5.7 and 19.1 points respectively. The open-weight version, named Qwen3.8-Flash-Next, was simultaneously published on Hugging Face and ModelScope.

The cost reduction is structural, not a subsidy. Four architectural innovations — Qwen Sparse Attention, a Gated Residual mechanism, N-gram Embedding that offloads 51 billion parameters to host memory, and a refined Muon optimizer — collectively reduce inference pressure on GPU high-bandwidth memory. Critically, the model achieves comparable performance to Qwen3.7-Plus at roughly one-ninth the training compute. The "Next" designation is a deliberate signal: this architecture serves as the preview for the forthcoming Qwen4 family, following the same sequencing Alibaba used when Qwen3-Next previewed the Qwen3.5 series. With hyperscaler inference capex now exceeding training capex industry-wide for the first time in 2026, architectural efficiency at the Flash tier has become the primary competitive battleground.


Zhipu's GLM-5.3 Flash Runs on 100,000+ Domestic Chips, Undercuts DeepSeek's Post-Hike Rates

Zhipu AI launched GLM-5.3 Flash on Wednesday at RMB 0.8 per million input tokens and RMB 2.8 per million output tokens — undercutting DeepSeek V4 Flash's post-hike off-peak output rate of RMB 4.5 per million tokens across virtually all standard call patterns. The 300-billion-parameter model, with 18 billion activated per forward pass, ran anonymously under the codename "Ox Alpha" on OpenRouter and OpenCode for five days prior to launch, accumulating more than 50 trillion tokens of free traffic and shattering both platforms' growth records. All inference runs on a cluster exceeding 100,000 domestic AI chips — likely including Huawei, Moore Threads, and Hygon hardware — with Zhipu claiming hardware efficiency and per-token cost at parity with Nvidia GPUs.

Semiconductor research firm SemiAnalysis flagged the deployment as a direct challenge to Nvidia's CUDA moat, noting that sustaining 100 trillion tokens per day on domestic silicon was previously assumed to require only the most well-resourced frontier labs. The timing is precise: DeepSeek's price hikes cut call volume on third-party platform OpenCode by half, creating a demand gap Zhipu is now priced to absorb. The model also marks Zhipu's return to native multimodal capability — supporting image and video input — broadening its enterprise addressable market beyond the coding workloads that defined its 2025 positioning. Whether domestic chip infrastructure can scale with the reliability and toolchain depth of Nvidia's ecosystem remains the critical open question.


MiniMax's ARR Hits $800M in August, but 17.9% Gross Margin Exposes a Dangerous Middle Ground

MiniMax reported H1 2026 revenue of US$116.6 million — a 283% year-on-year surge that already exceeds full-year 2025 revenue — and disclosed that annualized recurring revenue crossed US$800 million in August, against a sell-side median estimate of roughly US$600 million. Enterprise API and open-platform services drove the beat, surging 703% to US$73.9 million and now accounting for 63.4% of revenue, up from roughly 30% a year ago. Token consumption in July was 20 times the January level. The company's developer and enterprise customer base has grown to over 2 million, approximately ten times the year-end 2025 figure.

The revenue beat, however, masks a structural margin problem. Gross margin fell to 17.9% from approximately 30% in Q4 2025, as MiniMax finds itself caught between two unfavorable poles: it lacks the frontier model capability to command premium pricing, yet cannot match DeepSeek's cost-per-token floor. Its M3 flagship model — whose training corpus weighted native multimodal data at the expense of text, degrading coding benchmark scores at exactly the moment enterprise buyers were willing to pay a premium for code generation — has fallen materially behind peers on third-party benchmarks. Three near-term catalysts could shift the narrative: M3.1 (a coding-focused post-training patch, expected August-September), M3 Pro (scaling to approximately 2.7 trillion total parameters, September-October), and the H3 video model, which has accumulated over 24 million downloads since its late-July release. At roughly 17x ARR on its current US$13.5 billion implied valuation, the risk-reward has shifted — but execution on M3.1's coding benchmarks is the binary near-term test.


Li Auto's Q2 Gross Margin Halves to 11% as In-House Chip and Battery Bets Deepen

Li Auto reported Q2 2026 net loss of RMB 1.71 billion (US$237.5 million), reversing a RMB 1.10 billion profit in Q2 2025, with vehicle gross margin collapsing to 9.4% from 19.4% a year earlier. Deliveries of 98,330 units fell 11.5% year-on-year. Q3 delivery guidance of 95,000–100,000 units — midpoint 97,500 — fell 20% below Bloomberg consensus of approximately 121,900 units. Revenue guidance of RMB 26.6–28.0 billion trails consensus by approximately 17% at the midpoint. Sequential improvement was visible — gross margin recovered to 11.0% from Q1's 7.9%, and operating cash flow turned marginally positive — but the year-on-year deterioration is the sharpest in the company's public history.

The strategic context is critical. CEO Li Xiang confirmed that all Li Auto models will carry proprietary battery technology "within the next few months," completing full-stack powertrain vertical integration after the company's self-developed MACH M100 chip entered mass production in May. More than 50,000 chip-equipped Li L9 units were delivered by end of Q2. Li Xiang explicitly invoked Apple and Huawei as the integration model: "Batteries and chips are the most critical moats." With RMB 87.5 billion in liquidity and six consecutive quarters of sustained R&D spend near RMB 2.8 billion, Li Auto has the runway to absorb the transition trough. The Q4 volume ramp — anchored on the Li i9 launch in mid-September targeting the RMB 400,000-plus segment — will determine whether Q2 marks a trough or a plateau.


What to Watch Next

The next 60 days will be decisive across multiple fronts simultaneously. DeepSeek's second funding round close and subsequent financial disclosures will set the valuation anchor for China's AI IPO pipeline. MiniMax's M3.1 coding benchmark performance is the single most important near-term signal for whether the company's underperformance reflects a recoverable training-mix error or a structural capability gap. Alibaba's September-quarter earnings — specifically external cloud revenue growth against a 50%-plus threshold and AI lab operating loss trajectory — will determine whether JPMorgan's Overweight thesis holds. And Zhipu's ability to sustain paid call volume conversion after its 50-trillion-token free-traffic blitz will test whether domestic chip infrastructure can underpin a durable commercial model at scale.

Related Coverage:

Li Auto’s Q2 Margin Halves as In-House Tech Bet DeepensMiniMax’s ARR Tops $800M, but Margin Squeeze Tests Its AI Growth ModelZhipu Undercuts DeepSeek With GLM-5.3 Flash on 100,000+ Domestic ChipsAlibaba's Qwen3.8-Flash Rewrites the AI Cost Curve, Previews Qwen4 ArchitectureDeepSeek Posts 10x Revenue Surge, 82.9% API Margin as Valuation Nears $70BJPMorgan Says Alibaba’s $19B Selloff Overstates the Cost of Its $10B Share Sale

Subscribe to ChinaBiz Insider

Don’t miss out on the latest issues. Sign up now to get access to the library of members-only issues.
[email protected]
Subscribe