Xiaomi Slashes AI Model Pricing by Up to 99% in Industry Race

Xiaomi Slashes AI Model Pricing by Up to 99% in Industry Race

Xiaomi Technology announced a permanent price reduction of up to 99% for its MiMo-V2.5 series APIs on May 27, 2026, marking the second major Chinese AI model provider to implement aggressive pricing cuts following DeepSeek's similar move earlier this year.

The Beijing-based technology giant eliminated traditional pricing tiers based on context window length while optimizing its token-based billing structure. Under the new framework, developers receive 5 to 8 times more token usage at equivalent price points, fundamentally reshaping cost economics for enterprise AI deployments.

The restructuring reflects intensifying competition in China's large language model sector, where operational efficiency and razor-thin margins increasingly determine market positioning rather than pure technological differentiation.

Granular Pricing Reveals Strategic Intent

MiMo-V2.5-Pro's input cache hit pricing dropped to RMB 0.025 per million tokens (US$0.0035), representing a 98% reduction from the ≤256k specification's previous RMB 1.40 rate and a 99% cut versus the 256k-1M tier's RMB 2.80. Input cache miss pricing settled at RMB 3.000 per million tokens, down 57% from RMB 7.00 and 79% from the long-window RMB 14.00 rate. Output token costs decreased to RMB 6 per million, marking 71% and 86% reductions from previous RMB 21 and RMB 42 pricing respectively.

The standard MiMo-V2.5 variant experienced comparable compression. Input cache hit pricing now stands at RMB 0.020 per million tokens, down 96% from the ≤256k tier's RMB 0.56 and 98% from the extended window's RMB 1.12. Input cache miss pricing reached RMB 1.000, reflecting 64% and 82% reductions from RMB 2.80 and RMB 5.60 baselines. Output costs dropped to RMB 2 per million tokens, declining 86% and 93% from RMB 14 and RMB 28 respectively.

Notably, Xiaomi maintained existing pricing for its MiMo-V2-Pro and MiMo-V2-Omni premium models while discontinuing associated token plan packages, channeling developers toward the cost-optimized V2.5 series. The MiMo-V2.5-TTS voice synthesis model continues operating under a free-access policy.

Leadership Transition Drives Execution Velocity

The pricing overhaul stems from organizational changes initiated in November 2025, when Xiaomi recruited Luo Fuli, a 27-year-old AI engineer formerly with DeepSeek, to lead MiMo model development. CEO Lei Jun reportedly offered Luo an annual compensation package exceeding RMB 10 million (US$1.39 million) to assemble a research team averaging 25 years old, with over 60% holding degrees from Tsinghua or Peking universities.

Under Luo's direction, Xiaomi released MiMo-V2-Pro, MiMo-V2-Omni, and MiMo-V2-TTS foundation models in March 2026, subsequently iterating to V2.5 variants that integrate high-performance reasoning, lightweight interaction protocols, and voice synthesis capabilities. On May 26, Lei announced MiMo-V2.5-Pro achieved joint first-place rankings globally among open-source models in Artificial Analysis benchmarks for comprehensive intelligence and agent task performance, alongside committing RMB 60 billion (US$8.33 billion) to AI investments through 2029.

Market Dynamics Shift Toward Efficiency Competition

DeepSeek established the current pricing template when it permanently reduced DeepSeek-V4-Pro API costs by 75% following a promotional period ending May 31, 2026. Post-adjustment rates match Xiaomi's structure: RMB 0.025 per million tokens for input cache hits, RMB 3 for cache misses, and RMB 6 for outputs. DeepSeek-V4, released in late April 2026, features million-character context windows and dominates Chinese and open-source rankings for agent capabilities, world knowledge, and reasoning performance.

These pricing actions underscore bifurcation in China's AI model landscape. General-purpose providers including Alibaba Cloud's Qwen, ByteDance's Doubao, Xiaomi, and DeepSeek pursue volume through aggressive cost reduction, while enterprise-focused vendors such as Zhipu GLM and Tencent Hunyuan maintain or incrementally raise pricing for customized implementations.

According to AI.cc's 2026 AI API Infrastructure Report, enterprise-grade large language model token costs declined 67% year-over-year, with open-source models capturing 38% of corporate token consumption volumes. The data points to fundamental shifts in underlying algorithm optimization, inference architecture improvements, and compute cost compression rather than temporary promotional tactics.

The strategic calculus now centers on converting cost leadership into sustained developer ecosystem lock-in, as token pricing approaches commodity status and differentiation migrates to fine-tuning capabilities, vertical domain specialization, and integration friction reduction. Xiaomi's elimination of context window pricing tiers and token plan consolidation signals recognition that transparent, simplified billing structures reduce switching costs—a counterintuitive approach when competitive moats erode.

For international observers, the pricing trajectory indicates Chinese AI model providers increasingly compete on operational leverage rather than technological moonshots, potentially foreshadowing similar compression in Western markets where OpenAI, Anthropic, and Google maintain premium positioning but face mounting pressure from open-source alternatives and efficiency-driven challengers.

Related Coverage:

DeepSeek Slashes API Pricing by 97.5%, Triggering Agent Model Cost War

Subscribe to ChinaBiz Insider

Don’t miss out on the latest issues. Sign up now to get access to the library of members-only issues.
[email protected]
Subscribe