Xpeng and Li Auto Debate VLA: Who is Naked Running, Who is Gambling?
The race for autonomous driving supremacy in China has intensified as executives from two leading electric vehicle makers doubled down on Vision-Language-Action (VLA) models as the definitive path to full self-driving. This public reaffirmation underscores a strategic divergence in the industry, pitting VLA proponents against advocates of World Models, as Chinese automakers scramble to prove their technical prowess in an increasingly competitive market.
Lang Xianpeng, head of autonomous driving at Li issued a strong defense of the technology on social media, asserting that VLA represents the "best model solution for autonomous driving." His comments directly countered earlier skepticism from robotics experts, emphasizing that future embodied intelligence will hinge on comprehensive system capabilities rather than simplified architectures.
Simultaneously, He Xiaopeng, founder of Xpeng, escalated the commitment by revealing a high-stakes internal wager. He pledged that if Xpeng's domestic VLA technology does not match the performance of Tesla’s Full Self-Driving (FSD) V14.2 in Silicon Valley by August 30, 2026, the company's autonomous driving chief, Liu Xianming, would face the penalty of a "naked run" at the Golden Gate Bridge. Conversely, success would result in the construction of a stylized Chinese cafeteria in Silicon Valley.
These aggressive posturing and ambitious timelines signal to investors that Chinese EV startups are prioritizing software sovereignty to narrow the gap with global leaders like Tesla Inc. As the industry approaches a critical inflection point for high-level driver assistance systems, the debate highlights the urgent pressure on manufacturers to validate their R&D investments and secure consumer mindshare before the next wave of consolidation.
Doubling Down on End-to-End AI
The current controversy was ignited by comments from Wang Xingxing, founder of Unitree Robotics, who previously dismissed VLA models as a "relatively foolish architecture." In a lengthy rebuttal, Li Auto’s Lang argued that VLA allows vehicles to leverage general-world knowledge—akin to the methodology behind Generative Pre-trained Transformers (GPT)—to solve "long-tail" scenarios. This approach aims to give autonomous systems human-like social experience, such as distinguishing between a hitchhiker and a traffic officer using hand gestures.
Xpeng is pushing this trajectory further with its VLA 2.0, scheduled for pioneer user testing in December 2025. According to the company, the core evolution of VLA 2.0 involves removing traditional rule-based logic layers ("L"), generating action commands directly via implicit logic. He Xiaopeng claims this architectural shift will significantly reduce latency and increase the miles between necessary driver interventions in complex urban environments, such as "urban villages," by 13 times.
The Divide Between VLA and World Models
While Xpeng and Li Auto champion VLA, other major players have chosen divergent technical paths. Executives from Huawei have explicitly stated that they are bypassing VLA in favor of World Models (WA) as the ultimate solution. Similarly, William Li of NIO has pledged to position Nio’s World Model at the top of the industry.
However, the technical distinctions may be blurring as companies seek to optimize efficiency. Xpeng’s removal of explicit logic layers brings its VLA closer to the underlying principles of World Models, which focus on minimizing information loss and maximizing the efficiency of image tokenization. Concurrently, Li Auto is utilizing World Models in the cloud to generate data, run simulations, and conduct reinforcement training. This convergence suggests that regardless of the terminology—or whether the "Language" component is explicit—the industry is collectively moving toward systems that can intuitively understand and predict physical world interactions, similar to Yann LeCun’s concept of Joint Embedding Predictive Architecture (JEPA).
The Battle for Narrative and Survival
The intensifying rhetoric serves a dual purpose: technical validation and marketing survival. With Tesla preparing to re-assert its dominance with FSD V14, Chinese automakers are under immense pressure to demonstrate comparable capabilities to maintain their valuation premiums. He Xiaopeng’s candid admission that their current version has not yet reached the level of FSD V14.2 underscores the reality behind the "naked run" wager—it is a recognition of the gap that still exists.
For investors, these public bets and technical debates act as signals of corporate confidence and urgency. As the aggressive "word creation" phase of marketing fades in 2025, the focus is shifting towards delivering tangible user experiences. Securing the narrative leadership in autonomous driving is no longer just about innovation; it is about establishing a defensive moat against both domestic rivals and global tech giants as the sector matures.