Chinese Robotics Firm Unitree Releases Open-Source VLA Model for Humanoid Manipulation

Chinese Robotics Firm Unitree Releases Open-Source VLA Model for Humanoid Manipulation

Unitree Robotics has open-sourced a vision-language-action model designed specifically for humanoid robot operations, marking a shift in industry focus from semantic understanding to physical task execution as commercial deployment becomes increasingly critical.

The company released UnifoLM-VLA-0 on January 29, positioning it as a model that bridges the gap between visual comprehension and reliable physical interaction. Unlike traditional vision-language models that excel at identifying objects and suggesting actions, this system is optimized for task completion rates—a metric Unitree identifies as essential for real-world deployment.

The release comes as the robotics sector pivots toward practical applications beyond demonstrations. With humanoid robots approaching mass production, the industry faces mounting pressure to deliver systems capable of performing useful tasks in unstructured environments rather than controlled showcases.

Unitree's decision to open-source the model provides the robotics community with a framework explicitly designed for manipulation tasks, offering an alternative approach to companies developing proprietary systems.

From Semantic Understanding to Physical Execution

UnifoLM-VLA-0 builds on the open-source Qwen2.5-VL-7B framework, with modifications targeting robot manipulation scenarios. The architecture incorporates an action prediction head that generates continuous action sequences rather than single-step decisions, combined with forward and inverse dynamics constraints to establish relationships between actions and physical outcomes.

The training regimen emphasizes data quality over volume. Unitree utilized approximately 340 hours of real robot data, selected for task relevance rather than quantity. The training dataset spans multiple dimensions including 2D detection and segmentation, 3D object detection, spatial reasoning, task decomposition, and trajectory prediction.

This approach addresses a fundamental challenge in robotics: aligning semantic information with spatial geometry. Traditional vision-language models can identify what needs to be done but lack the physical common sense required to execute tasks reliably in three-dimensional space.

Benchmark Performance and Spatial Capabilities

In spatial understanding benchmarks including ERQA, RoboSpatial, and Where2Place, UnifoLM-VLA-0 demonstrated substantial improvements over its Qwen2.5-VL-7B foundation. Operating in "no thinking" mode, the model achieved performance comparable to Gemini-Robotics-ER 1.5.

The LIBERO simulation benchmark revealed strong task generalization. UnifoLM-VLA-0 achieved near-perfect scores across Spatial, Object, and Goal subsets, maintaining high completion rates even in Long scenarios involving extended task sequences. The overall average score reached 98.7 across four scenario categories.

This single-policy generalization represents a departure from traditional approaches requiring separate training for each task type, potentially reducing deployment complexity and operational costs for commercial applications.

Real-World Validation on G1 Platform

Unitree conducted physical testing on its G1 humanoid robot across 12 manipulation tasks using a unified policy network for end-to-end execution. Several tests incorporated external disturbances to evaluate system robustness.

In a towel-folding task, the G1 robot flattened and folded the towel through multiple steps. When an operator unfolded the towel mid-task, the system recognized the state change and repeated the folding sequence. During a stationery organization task, the robot successfully placed items in appropriate positions despite objects being removed twice during execution.

A fruit-sorting demonstration required placing colored items into matching containers. After positioning an avocado, the red container's location shifted. The robot adapted to the real-time change and correctly placed a watermelon in the relocated container.

The tests confirmed task completion across all categories without task-specific parameter adjustments, with the system maintaining execution stability under external perturbations.

Industry Implications

The open-source release addresses practical concerns as humanoid robotics transitions from research demonstrations to commercial deployment. While vision and language capabilities have matured, generating reliable executable actions in physical environments remains a bottleneck.

Unitree's approach prioritizes task success rates over semantic sophistication, reflecting broader industry recognition that commercial viability depends on robots delivering concrete value beyond performance capabilities. As production scaling accelerates, the focus has shifted from what robots can potentially do to what they reliably accomplish in operational settings.

The model's project page and source code are available through the company's GitHub repository, providing researchers and developers access to the complete framework for humanoid manipulation tasks.

Subscribe to ChinaBiz Insider

Don’t miss out on the latest issues. Sign up now to get access to the library of members-only issues.
[email protected]
Subscribe