XPeng's Second-Generation VLA to Launch March 2 With Volkswagen as Launch Customer, Signaling China's Growing AI Driving Ambitions

XPeng's Second-Generation VLA to Launch March 2 With Volkswagen as Launch Customer, Signaling China's Growing AI Driving Ambitions

XPeng is set to unveil its second-generation Vision-Language-Action model on March 2, positioning the technology not merely as an incremental upgrade to its autonomous driving stack but as what the company calls the world's first mass-production-grade physical world large model — a distinction that carries significant implications for the global race toward full autonomy.

Volkswagen AG will serve as the launch customer for the new model, a development that underscores a broader shift in the competitive dynamics of intelligent driving technology. The pairing of a legacy European automaker with a Chinese AI-native EV company as a technology recipient marks a notable reversal of the traditional direction of automotive technology transfer.

XPeng Chairman He Xiaopeng announced the launch date via social media on February 26, confirming that the 2026 XPeng X9 pure electric version will be the first production vehicle to carry the system. He Xiaopeng has also stated that the second-generation VLA will be made open-source for global commercial partners, a move that could accelerate adoption and reshape industry standards around physical AI.

A Structural Break From Conventional Autonomous Driving Architecture

The core technical claim behind the second-generation VLA is architectural rather than incremental. Conventional intelligent driving systems rely on a three-stage pipeline — vision, language, and action — in which visual inputs are first translated into linguistic representations before being converted into driving commands. XPeng's new model eliminates the intermediate language layer entirely, enabling direct end-to-end generation from visual signals to action instructions.

The practical significance of removing this translation step lies in reduced information loss and faster inference. By compressing a three-stage process into a unified model, the system is designed to respond more rapidly to dynamic real-world conditions — a critical requirement as autonomous driving systems approach higher levels of operational complexity.

This architectural shift is what XPeng is using to justify the label "physical world large model," distinguishing the VLA from software systems that operate primarily in digital or linguistic domains.

Scale of Training Data and Compute Infrastructure

The second-generation VLA was trained on close to 100 million video clips of real-world driving footage, a dataset that XPeng characterizes as equivalent to the extreme-scenario driving experience a human driver would accumulate over 65,000 years. The scale of the training corpus is central to the model's claims of robustness across edge cases and rare road conditions.

On the compute side, the model was developed on a cloud infrastructure comprising 30,000 accelerator cards, with a foundation model of 72 billion parameters. For vehicle-side inference, XPeng relies on its proprietary Turing AI chip, which delivers 2,250 TOPS of onboard compute. The company describes the resulting system as a fully integrated optimization across chip, operator, and model layers — a full-stack approach that gives XPeng control over each layer of the AI driving stack.

Volkswagen's decision to adopt the Turing AI chip as a designated component further validates the commercial viability of XPeng's silicon strategy, suggesting that the chip has met the qualification standards of a major global OEM.

A Platform Designed for the L4 Era and Beyond

XPeng is framing the second-generation VLA not as a feature for current production vehicles but as the foundational layer for a future in which full autonomy is commercially viable. He Xiaopeng stated in his 2026 new year letter to employees that the inflection point for fully autonomous driving is now clearly visible — a more definitive claim than the company has made in prior years.

The model's architecture is designed for cross-domain deployment. Beyond passenger vehicles, XPeng intends the same underlying physical world model to drive its Robotaxi service, its humanoid robot IRON, and its flying car program. This multi-platform strategy reflects a broader industry thesis: that a sufficiently capable model trained on physical-world interaction data can generalize across robotic systems, not just automobiles.

The commercial rollout timeline calls for a full over-the-air push to XPeng Ultra vehicle owners during the first quarter of 2026.

Volkswagen Partnership Tests the Limits of China's Technology Export Potential

The strategic weight of the Volkswagen relationship extends beyond a single customer win. It represents a test case for whether Chinese autonomous driving technology can be exported and integrated into global OEM platforms at scale — a question that has significant implications for both the competitive positioning of Chinese AI companies and the sourcing strategies of established automakers navigating the transition to intelligent vehicles.

He Xiaopeng's decision to open-source the second-generation VLA for global partners amplifies this dynamic. If the model gains traction as an open platform, XPeng could shift from being a vehicle manufacturer competing in a crowded EV market to becoming an infrastructure provider for physical AI — a structurally more defensible and higher-margin position.

Whether that ambition is achievable will depend on regulatory acceptance in key markets, the model's real-world performance relative to competing systems, and the pace at which global OEMs are willing to cede control over core driving intelligence to external technology providers.

Subscribe to ChinaBiz Insider

Don’t miss out on the latest issues. Sign up now to get access to the library of members-only issues.
[email protected]
Subscribe