Li Auto Unveils MindVLA-o1, Pushing In-Car Autonomy Toward a Shared “Physical AI” Stack
Li Auto used NVIDIA GTC 2026 to position its next-generation autonomy foundation model, MindVLA-o1, as more than a driver-assist upgrade, framing it as a unified “vision-language-action” system designed to scale from vehicles to robots.
At the conference on March 17, 2026, Li Auto base model lead Kun Zhan introduced MindVLA-o1 and said that unifying vision, language and action inside one model turns autonomy into an early application of a broader “physical-world general agent,” with automotive deployment serving as the starting point.
The announcement matters for investors because it signals Li Auto’s strategy to compound product capability through a reusable AI stack, and to lower iteration cost via simulation and hardware-aware design, two levers that can accelerate feature rollouts and reduce the marginal cost of improving autonomy.
Li Auto also highlighted adoption metrics for its prior VLA driver model, pointing to scale effects from real-world usage data that the company argues will support faster training and validation of the next model generation.
Scaling Product Usage Builds a Data Flywheel
Li Auto said it has iterated its assisted-driving architecture since launching in-house R&D in 2021, with 2024 marking a shift as an “end-to-end + VLM” dual-system architecture entered mass production and delivery, aiming to improve cross-scenario understanding.
In 2025, the company unified spatial understanding, language understanding and action decision-making into a single framework, built around VLA, world models and reinforcement learning. Li Auto said the VLA driver model was delivered with Li i8 in August and rolled out to AD Max users in September.
By the end of 2025, Li Auto said monthly usage of the VLA driver model reached 80%, with cumulative VLA instruction usage totaling 12.254 million. During the Lunar New Year period, Li Auto said assisted-driving mileage reached 250 million kilometers and VLA instructions were used 1.303 million times, underscoring how real-world usage can feed model iteration.
Packaging Five Innovations Reframes Autonomy as a Foundation Model
Li Auto said MindVLA-o1 is built on a native multimodal MoE Transformer and combines five technical innovations: 3D spatial understanding, multimodal thinking, unified action generation, closed-loop reinforcement learning and hardware-software co-design.
For perception, the company described a vision-centric 3D ViT encoder, using LiDAR point clouds as 3D geometric prompts to guide spatial understanding, and adding a feedforward 3D representation that separates static environment and dynamic objects, using next-state prediction as a self-supervised signal.
For reasoning, Li Auto said it uses a predictive latent world model trained in three phases, aiming to simulate how scenes evolve seconds ahead in latent space, then align world modeling, multimodal reasoning and driving behavior. The company labeled this capability “Generative Multimodal Thinking.”
For action, Li Auto said a VLA-MoE architecture includes a dedicated Action Expert to generate driving trajectories from 3D scene features, navigation goals and driver instructions, supported by parallel decoding for real-time output and discrete diffusion to iteratively refine trajectories under vehicle dynamics constraints.
Cutting Training Cost and Deployment Time Targets Faster Iteration
Li Auto said MindVLA-o1 incorporates a closed-loop reinforcement learning framework that learns from both real data and exploration inside a world simulator, enabled by upgrading reconstruction to feed-forward scene rebuilding for instant, large-scale, high-fidelity scenarios.
The company said it developed a unified 3D Gaussian Splatting rendering engine and distributed training framework, improving rendering speed by nearly 2x and reducing overall training cost by about 75%, aiming to make large-scale RL training more economical and frequent.
On deployment, Li Auto introduced what it called a hardware-software co-design “law” for on-device large models, combining model-structure and validation-loss modeling with a Roofline-based view of compute and memory bandwidth limits. The team evaluated nearly 2,000 architecture configurations and validated them on Nvidia’s Orin and Thor platforms, finding a Pareto frontier between accuracy and inference latency and shrinking architecture exploration from months to days.
Extending the Stack Beyond Cars Raises the Stakes for Platform Strategy
Li Auto positioned MindVLA-o1 as one module in a broader “digital brain” framework that also includes MindData (a unified VLA data engine for collection, cleaning and auto-labeling), MindSim (a controllable multimodal world model for generating complex scenarios), and RL Infra (reinforcement learning infrastructure using reward modeling and policy learning).
By pitching cars as “the largest robot,” Li Auto is arguing that the same core model and tooling can generalize to other physical systems, a strategy that, if executed, would shift investor focus from single-product features toward the reusability and efficiency of the company’s AI pipeline across multiple form factors.
Related Coverage:
Li Auto's Margin Slump Triggers Strategic Pivot Toward AI Chips and Robotics