1T Model, 1,000 Tokens/s, 8 GPUs: Xiaomi Redefines Inference Limits

1T Model, 1,000 Tokens/s, 8 GPUs: Xiaomi Redefines Inference Limits

Xiaomi and inference specialist TileRT have achieved what the industry has long treated as a hardware problem — cracking the 1,000 tokens-per-second threshold on a one-trillion-parameter large language model using a single standard eight-GPU node, a result that directly challenges the commercial logic of custom-silicon AI inference.

The joint announcement, made June 9, 2026, centers on the UltraSpeed mode of Xiaomi MiMo-V2.5-Pro, the company's flagship reasoning model. The benchmark that crystallizes the achievement: a complex AI operations dashboard — full HTML, CSS, JavaScript, canvas animations, and real-time data visualizations — generated in 13 seconds flat. The same task on the standard MiMo-V2.5-Pro took 6 minutes and 15 seconds, a 28x speed differential that transforms the model from a batch-processing tool into something approaching real-time infrastructure.

Xiaomi founder Lei Jun amplified the announcement on Weibo, framing the milestone as a "10x speed, 3x price" proposition — a deliberate signal to enterprise developers weighing latency costs against inference budgets.


Commoditizing Speed: Why Eight GPUs Reframe the Competitive Map

The strategic significance of this announcement lies not in the raw speed figure but in the hardware constraint under which it was achieved. Rivals pursuing sub-second inference at trillion-parameter scale have largely bet on custom silicon: Cerebras Systems deploys wafer-scale integration, while Groq relies on on-chip SRAM architecture that eliminates memory bandwidth bottlenecks by design. Both approaches demand specialized procurement pipelines and carry substantial capital expenditure premiums — barriers that effectively price out mid-tier cloud operators and enterprise AI teams.

Xiaomi and TileRT instead attacked the same bottleneck through a co-designed software stack, combining three interlocking techniques:

FP4 quantization compresses model weights using the MXFP4 standard, applied selectively to the Mixture-of-Experts (MoE) expert layers of MiMo-V2.5-Pro while preserving full precision elsewhere. Quantization-Aware Training (QAT) ensures benchmark parity with the uncompressed model, addressing the traditional accuracy-versus-efficiency tradeoff that has made aggressive quantization commercially risky.

DFlash speculative decoding replaces the conventional draft-model-plus-verifier pipeline with block-level masked parallel prediction. Rather than generating candidate tokens sequentially, the draft model fills an entire masked block in a single forward pass, then submits the batch for simultaneous verification. In coding tasks — the highest-value commercial use case — the team reports an average accepted token length of 6.30 per verification cycle, with peak samples reaching 7.14 out of every 8 drafted tokens. That acceptance rate directly converts into throughput: fewer verification cycles per output token means fewer round-trips through the full trillion-parameter model.

TileRT's resident kernel engine eliminates the operator-boundary gaps that become the dominant latency source at 1,000 tokens/second. Rather than launching discrete compute kernels per operation, TileRT maintains a persistent execution pipeline inside the GPU, overlapping data movement and computation at the warp level. The result is what the team describes as "microsecond-level hardware convergence" — execution pressure that closes within the physical limits of the GPU rather than spilling into scheduling overhead.


Pricing Architecture Reveals a Two-Tier Market Strategy

The commercial structure of the UltraSpeed launch deserves scrutiny from enterprise procurement teams. Following Xiaomi's May 27, 2026 announcement of permanent price reductions across the MiMo-V2.5 series, the base MiMo-V2.5-Pro API now prices at RMB 0.025 per million tokens (cache hit input), RMB 3 per million tokens (cache miss input), and RMB 6 per million tokens (output).

UltraSpeed carries a 3x multiplier across all tiers, implying output pricing of RMB 18 per million tokens (approximately US$2.50 at current reference rates). For latency-insensitive workloads — document summarization, offline data enrichment, batch classification — the standard tier remains cost-optimal. For the use cases Xiaomi explicitly targets with UltraSpeed — high-frequency trading signal generation, real-time fraud detection, surgical assistance, and agentic coding pipelines — the 3x premium against a 10x speed gain produces a favorable cost-per-second-of-latency calculation that will appeal to operators where response time carries direct revenue or risk implications.

The access model, however, signals supply constraints. The initial rollout runs June 9–23, 2026, operates on an application basis, and does not support Token Plan subscriptions — only direct API calls. Approved users receive two weeks of complimentary Chat access. This controlled rollout structure is consistent with GPU cluster capacity management rather than deliberate market rationing, suggesting that scaling UltraSpeed to general availability remains an infrastructure challenge.


Open-Source Release Accelerates Ecosystem Adoption

Xiaomi has published the MiMo-V2.5-Pro-FP4-DFlash checkpoint to HuggingFace, including both FP4 quantized weights and DFlash model parameters. This move follows the broader industry pattern of using open-weight releases to build developer ecosystems around proprietary API services — a strategy that DeepSeek demonstrated effectively in early 2025, driving API adoption by enabling local experimentation before commercial deployment.

For TileRT, the collaboration with Xiaomi represents a significant reference deployment. The inference infrastructure firm previously set what was then the public commercial API speed record on May 22, 2026, in a joint optimization with Zhipu AI (智谱AI) that pushed GLM-5.1's high-speed API to 400 tokens/second. The jump from 400 to 1,000 tokens/second in under three weeks, on a model 2.5x larger in parameter count, indicates rapid iteration in TileRT's compilation and kernel optimization stack.


Limitations Constrain Near-Term Commercial Scope

The technical disclosure contains one notable caveat that investors and enterprise evaluators should weight carefully. DFlash's high acceptance rates — the mechanism that drives the throughput gains — are concentrated in structured, predictable output tasks: coding, templated generation, and agent workflows where token sequences follow learnable patterns. In open-ended conversational contexts, where output entropy is higher, acceptance rates remain "not high" by the team's own characterization, with optimization ongoing.

This constraint matters because the largest volume tier of LLM API consumption globally remains general-purpose chat and document interaction, not agentic coding. Until DFlash acceptance rates generalize across higher-entropy tasks, the 1,000 tokens/second figure applies to a subset of enterprise workloads rather than the full API surface. The team's acknowledgment of this gap is technically honest; it also defines the next performance frontier that will determine whether UltraSpeed transitions from a specialized premium tier to a broadly deployable infrastructure standard.

Related Coverage:

Xiaomi Slashes AI Model Pricing by Up to 99% in Industry Race

Subscribe to ChinaBiz Insider

Don’t miss out on the latest issues. Sign up now to get access to the library of members-only issues.
[email protected]
Subscribe