Xiaomi Slashes AI Agent Costs with MiMo-V2.5 Launch
Xiaomi has escalated China's artificial intelligence price war with the sudden launch of its MiMo-V2.5 large language model series, aggressively targeting token efficiency and complex agent automation just days before the anticipated release of DeepSeek V4.
Led by former DeepSeek core developer Luo Fuli, the release arrives a mere 36 days after the previous iteration. The rapid deployment underscores the brutal iteration cycle in the 2026 Chinese AI sector, where tech giants are pivoting from raw parameter size to operational efficiency and system-level integration.
Markets are assessing the strategic timing of the launch. By restructuring its API pricing and demonstrating autonomous capabilities in software engineering and hardware design, Xiaomi is positioning its AI not merely as a chatbot, but as the underlying operating system for its broader consumer electronics and electric vehicle ecosystem.
Slashing Token Costs Intensifies AI Margin Squeeze
The core value proposition of MiMo-V2.5 lies in its computational efficiency. According to benchmark data on the ClawEval agent framework, the flagship MiMo-V2.5-Pro requires 42% fewer tokens to achieve identical scores compared to Kimi K2.6, recently released by Chinese competitor Moonshot AI. Against Meta's Muse Spark, the standard V2.5 model demonstrated a 50% reduction in token consumption.
To capture enterprise developers, Xiaomi overhauled its subscription-based MiMo Token Plan. The company eliminated its previous 4x credit multiplier, unified billing for 256k and 1M context windows, and introduced an off-peak rate structure offering a 20% discount between midnight and 08:00. Annual subscribers receive discounts up to RMB 948 (US$137.39), a direct mechanism to lock in developers amidst intense domestic API price dumping.
Automating Complex Engineering Workflows
Xiaomi detailed the model's capacity for autonomous, long-horizon tasks, moving beyond simple code generation. Internal metrics show MiMo-V2.5-Pro can sustain logical consistency across nearly 1,000 tool calls in a single session.
In a demonstration of its engineering utility, the model autonomously constructed a complete SysY compiler in Rust—a task involving lexical analysis, Abstract Syntax Tree (AST) generation, and RISC-V assembly backend. The process required 4.3 hours and 672 tool calls, achieving a perfect score of 233 on a Peking University benchmark. Furthermore, the model executed an analog circuit Electronic Design Automation (EDA) task, designing a Low Dropout Regulator (LDO) on a TSMC 180nm CMOS process in approximately one hour—a workflow that typically takes experienced engineers several days.
Accelerating Hardware-Ecosystem Integration
The aggressive push into agentic AI serves a broader corporate strategy. In March, the MiMo-V2-Pro appeared anonymously on the OpenRouter platform as "Hunter Alpha," temporarily mistaken by developers for DeepSeek V4. Now officially branded, the omnimodal V2.5 model integrates image, audio, and video processing natively.
For Xiaomi, deploying a lightweight, high-efficiency model is less about selling API access and more about embedding system-level native agents across its "Human x Car x Home" hardware ecosystem. By lowering inference costs while maintaining high reasoning capabilities, the company clears the technical bottleneck for deploying persistent, always-on AI assistants across millions of edge devices and smart vehicles.
Related Coverage:
Xiaomi Claims Viral “Hunter Alpha” Models as MiMo V2 Trio, Pressuring China’s Agent AI Pricing