RobotEra Founder: Learning from Humans is 'Shortest Path' to General-Purpose Robots
In a keynote speech at the main forum of the 2025 World Robot Conference (WRC), Chen Jianyu, founder of Robotera, outlined the company's strategic approach to creating advanced automatons. Titled "Building General-Purpose Humanoid Robots," the presentation detailed a vision centered on integrating a "general brain" with a "general-purpose body" to unlock the next wave of AI-driven productivity.
Chen asserted that "learning from humans" represents the most direct route to developing truly versatile humanoid robots, which he positions as the next major frontier for artificial intelligence after its integration into personal computers and smartphones. His address provided a comprehensive overview of Robotera's rationale, development pathway, and recent achievements, including the delivery of over 300 units to major technology clients, offering a significant look into the strategy of a prominent player in the competitive global robotics landscape.
Below is the full transcript of Chen Jianyu's speech.
General Robots are the Next Frontier for AI, Set to Revolutionize Productivity and Services
"We recently launched two full-sized humanoid robots—one bipedal and one wheeled. Our humanoid robots can not only perform dynamic movements like high-impact 360-degree rotational jumps and breakdancing, but they are also capable of handling a wide variety of general-purpose operational tasks, such as logistics sorting, folding clothes, carrying objects, scanning codes, and tightening screws."
"We believe that general-purpose robots are definitively the next trend for AI. We can see that AI has already permeated various terminals like computers and mobile phones. Now, it is transitioning from 'thinking' to 'acting.' Intelligent vehicles are one example of this shift. Next, because of their more powerful and versatile mobility and manipulation capabilities, robots are poised to revolutionize productivity and service capabilities across society."
Traditional Hardware/Software Models Are Not Generalizable and Lead to Commercial Traps
"Why are we building such a general-purpose humanoid robot? We believe that relying on traditional robotic hardware and software solutions makes it difficult to achieve true versatility. Although there is already a wide variety of robots, their numbers are actually quite small compared to the major terminal products I just showed. This is because a separate, independent system must be built for each specific scenario. We argue that such an accumulation of hardware cannot foster the ultimate evolution of intelligence. This specialized systems approach ultimately leads to a commercial trap, preventing us from scaling robots effectively. This is why the entire field of robotics, despite over half a century of development, has yet to produce a true industry giant."
General-Purpose Humanoid Robot = General Brain + General-Purpose Body
Learning from Humans is the Shortest Path to General-Purpose Robots
"So, how do we build a general-purpose robot? We believe the shortest path is to learn directly from humans, as humans are the only general-purpose embodied intelligence that exists in the real world. Why have our language models been so successful? It is precisely because they mirror the process of human language acquisition, learning from vast amounts of human-generated text."
"Robotics involves a much broader range of dimensions. Therefore, we need to build a 'general brain' for the robot that not only possesses language capabilities but can also control hands and legs to interact with the physical world. Simultaneously, we must construct a matching 'general-purpose body' for it."
The General Brain (ERA-42): End-to-End Models as the Key to General-Purpose Robotics
"First, let's talk about the brain of a general-purpose robot. We have released a general brain model called ERA-42. It is an end-to-end embodied model that integrates visual perception, behavior understanding, planning, and execution into a single unified framework."
"Why create such an end-to-end model? Our inspiration comes from language models. Within just a few months of their emergence, language models completely disrupted the entire field of NLP (Natural Language Processing). The NLP field had previously constructed many different models and a multitude of algorithms to solve various tasks, but they were ultimately overturned by the elegant architecture of the Transformer, which demonstrated superior performance across the board. Therefore, we believe that robotics must follow a similar path to achieve the general-purpose model we ultimately desire."
"However, such a model still faces the challenge of controlling a general-purpose humanoid body with a single model. We have been working diligently to overcome this and have made some progress. We are now able to use a single model to control high-degree-of-freedom (DoF) robot bodies and achieve excellent performance with relatively limited training data."
Continuous Breakthroughs in Embodied Model Research Paradigms are Needed to Overcome Bottlenecks
"Underpinning this progress is our continuous effort to break through the research paradigms of embodied models. We believe the biggest bottleneck currently lies in the paradigm of the final embodied model itself. We must constantly break through and iterate on this paradigm to overcome existing limitations. We divide the development process of embodied models into four stages, which also represent the four exploratory stages of Robotera."
"Stage One: We explored how to introduce language models and visual-language models, which possess human-like cognitive abilities, into embodied intelligence. However, at this stage, the cognitive model and the action model were still separate entities. This was the common approach around 2023, shortly after ChatGPT was released."
"Stage Two: The current mainstream approach involves models based on a fast-slow system, such as π₀ and Helix. We call this 'Real-time Action, Deep Thinking'—combining the deep reasoning capabilities of language models with the real-time execution of actions into an end-to-end model. Although it is a fast-slow system, it is trained end-to-end. We began exploring this early on and published related papers in the middle of last year."
"Stage Three: This is represented by generative models like Sora. Why is this important? Robots have concrete interactions with the physical world, whereas language models remain at an abstract level of spatial understanding. Generative models like Sora are actually capable of capturing the paradigms of very fine-grained physical interactions."
"Crucially, these models can also learn the laws and knowledge of the physical world from vast amounts of unlabeled internet video data—what we call a 'world model.' This method indirectly addresses the data scarcity bottleneck, as it allows for self-supervised learning directly from massive amounts of unlabeled internet video."
"Stage Four: The reinforcement learning paradigm, represented by models like DeepSeek. It garnered significant attention because its R1 model utilized reinforcement learning. Previous VLA (Vision-Language-Action) embodied models were primarily based on imitation learning from human demonstrations. This approach has two main problems: first, the model cannot surpass the capabilities shown in the demonstration. Second, its performance on specific physical tasks is often suboptimal. We have also conducted explorations in this area, using reinforcement learning to train our VLA models, which has ultimately improved their success rates and overall effectiveness."
ERA-42 Pre-training: The "Open-Book Exam" for Learning by Watching
"In summary, the initial phase is the pre-training stage, which we call the 'open-book exam,' or 'learning by watching.' This is similar to a child who spends their first few years observing the world without performing many concrete actions. Our pre-training process is analogous. This stage not only incorporates various types of robot data but also includes massive amounts of unlabeled internet video data, creating a pre-trained model integrated with a world model. This model can achieve zero-shot generation of execution policies, and these policies can be visualized as high-definition videos, allowing the robot to foresee and simulate its actions in entirely new scenarios and tasks."
ERA-42 Real-World Fine-Tuning: "Mastering through Practice"
"With this pre-trained foundation, the model can be considered to possess general common sense about the world. The next step is concrete practice and optimization. This requires the model to collect real data in the physical world using its specific robot body, and then fine-tune itself based on that data."
"Because of the initial 'open-book exam' phase, we only need a very small amount of real-world robot data during the second stage to significantly improve task accuracy. This paradigm also effectively solves our data bottleneck problem."
ERA-42 Breaks the Data Bottleneck, Ensuring a Rich Learning Source
"This is a diagram of the Robot Data Pyramid that I've drawn. At the very top is real-world robot data, which is of the highest quality but extremely limited in quantity. For comparison, I've shown the volume of text or video data used for GPT-4 and Sora. The amount of real-world robot data is minuscule in comparison. Relying on this data alone, it would be very difficult to achieve the level of generalization we have already reached."
"Therefore, we introduced the two lower layers of the pyramid. One is human behavior data. The widespread development and gradual adoption of VR and smart glasses now allow us to efficiently collect first-person human behavior data at a cost far lower than collecting real-world robot data. At the base is the even more massive pool of internet data, which includes human actions (first-person, third-person, and multi-person interactions), natural phenomena, animal activities, and more. Essentially, the world model can learn from everything that happens on Earth. With this data architecture, as our models iterate, the amount of real-world robot data we require has been substantially reduced."
"Simultaneously, we have been continuously increasing the difficulty of controlling the robot body, conducting cross-task and cross-body learning. We began our experiments last year on a single robotic arm, then progressively scaled up to a 7-axis arm with a five-fingered dexterous hand, enabling our model to directly control each finger's movement end-to-end. Subsequently, we transferred this to a dual-arm humanoid robot and then to even more complete forms."
The General-Purpose Body: The Humanoid Form as the Ultimate General-Purpose Morphology
"The second part concerns the general-purpose body module. The keywords here are 'generalization,' 'modularization,' and 'full-sized humanoid.' Why are we building a humanoid robot? Because our human environment is built by humans, for humans. We believe the ultimate, most general-purpose form is the humanoid form. But building humanoid robots is not just the end goal but also the means—by making humanoid robots, we can collect more data at a lower cost. Furthermore, the first-person human behavior data and internet data I just mentioned can be more effectively transferred to our humanoid robot bodies."
Hardware Generalization and Modularization Enables Adaptation to Different Scenarios
"To enable our robot hardware to better adapt to various scenarios, we have adopted a strategy of hardware generalization and modularization. You can see that our modularity is multi-layered. At the top level is the whole-robot body, where we have the Robotera L7 for industrial applications and the Robotera Q5 for the service industry. Both are built upon the same set of joint modules and dexterous hands. The dexterous hands are also composed of smaller joint modules. These modules, in turn, contain core components like motors, reducers, and drivers, all of which are developed in-house. By developing our software and hardware ourselves, we ensure that the hardware is better adapted to the software, allowing them to evolve in synergy. This is a practical application of 'software-defined hardware'."
The "Model-Body-Data" Flywheel for AI Evolution in the Physical World
"In conclusion, the combination of a general brain and a general-purpose body allows us to establish a paradigm for building general-purpose humanoid robots. By integrating this with specific scenarios and data, we create an AI flywheel for the physical world. This means we build a unified model at the top level, which can universally empower various humanoid robot bodies (including dexterous hands). Different bodies can then adapt to different scenarios, and the applications in these scenarios generate feedback data, creating a closed-loop flywheel of continuous iteration and evolution."
"We have already achieved promising results with our physical AI flywheel. We were recognized by NVIDIA as one of the top 14 humanoid robot companies globally and were also included in Morgan Stanley's 2025 report on the humanoid robot industry as one of the top 16 worldwide. As of July this year, our product deliveries have surpassed 300 units. We have garnered interest from leading global technology giants; nine of the world's top ten technology companies by market capitalization are now our clients."