At Shanghai's World Artificial Intelligence Conference, China's robotics leaders confronted a paradox that echoes across many technological frontiers: the body has been built, but the mind remains starved. The machines are capable, yet the information needed to make them wise — drawn from the messy, varied, unpredictable physical world — exists in quantities far too small to close the gap between mechanical promise and genuine intelligence. It is a reminder that in the age of AI, the scarcest resource is not silicon or steel, but meaningful experience.
Chinese robotics firms cite data scarcity, weak AI models as key obstacles
How do you unlock these scenarios and replicate them at scale?
So when Wang talks about linking hardware, data, models, and real-world scenarios into a closed loop, what does that actually look like in practice?
It means you can't design a robot's arm in isolation and then hope the AI catches up later. The shape of the gripper, the way you collect video of a human hand grasping something, and the neural network learning from that video all have to be designed with each other in mind. Right now they're separate projects.
And the data scarcity—is that just a matter of time and money, or is there something structurally harder about it?
It's structural. Language models can be trained on any text that exists anywhere on the internet. But robot data has to come from robots actually operating in real environments. You can't just scrape it. You have to deploy machines, run them, record what happens, and do that thousands of times over in different scenarios.
So the bottleneck isn't that the AI models are bad, necessarily?
Not exactly. The models are constrained by what they're fed. You could have a brilliant architecture, but if it's only trained on data from five warehouses and three factories, it won't generalize to a sixth warehouse with different lighting or a different layout.
What would it take to break through this?
Massive deployment. Companies would need to put robots into far more diverse real-world situations and be willing to share or pool the data they collect. Right now everyone's protecting their own datasets like trade secrets.
El Pulso
- China's robotics firms arrive at a critical inflection point: their hardware is competitive, but without sufficient real-world training data, their robots cannot meaningfully improve.
- The bottleneck is structural — hardware, data collection, AI models, and deployment environments must evolve together, yet they remain disconnected silos that fail to reinforce one another.
- Multi-modal physical-world data is catastrophically scarce compared to the vast text corpora that powered the large language model revolution, leaving humanoid robot development without a comparable foundation.
- Industry leaders are pushing to scale real-world robot deployments and diversify operational scenarios, betting that broader exposure to unpredictable environments is the only path to generating the data they need.
- Until a self-reinforcing cycle of deployment, data generation, and model improvement is established, China's robotics sector risks stalling just as its mechanical ambitions reach their peak.
At Shanghai's World Artificial Intelligence Conference, China's robotics leaders confronted a paradox that echoes across many technological frontiers: the body has been built, but the mind remains starved. The machines are capable, yet the information needed to make them wise — drawn from the messy, varied, unpredictable physical world — exists in quantities far too small to close the gap between mechanical promise and genuine intelligence. It is a reminder that in the age of AI, the scarcest resource is not silicon or steel, but meaningful experience.
At last week's World Artificial Intelligence Conference in Shanghai, China's robotics leaders gathered around a shared frustration: the hardware is ready, but the intelligence is not.
Wang Xiaogang, co-founder of SenseTime and chair of its robotics division Ace Robotics, framed the challenge as architectural rather than technical. Advancing embodied AI requires four elements — the physical machine, the data it generates, the models that learn from that data, and the real-world environments where robots must perform — to function as a single, self-improving system. At present, those elements remain disconnected, unable to reinforce one another in the cycle that would drive genuine progress.
The data problem cuts deepest. Most training information comes from recording human demonstrations, but the way a robot's body is designed, the way that data is captured, and the physical structure of the machine must all be developed in concert. Companies are also deploying robots in environments that are too narrow and controlled to generate the volume and variety of real-world data needed. Wang's central question was blunt: how do you scale the situations where robots actually learn?
Yao Maoqing of AgiBot, a Shanghai humanoid robot maker, pointed to a parallel constraint. The multi-modal data describing how the physical world looks, feels, and behaves is nowhere near as abundant as the billions of text samples that trained large language models. This scarcity is blocking the development of world models — AI systems that would allow next-generation robots to understand their surroundings, manipulate objects, and adapt to the unexpected.
The irony is not lost on the industry. China has made genuine strides in mechanical engineering and hardware design. What stands between its robots and the next level of capability is not metal or motors — it is the depth and scale of lived, physical experience that only broad real-world deployment can provide.
At the World Artificial Intelligence Conference in Shanghai last week, the conversation among China's robotics leaders kept circling back to the same frustration: they have the hardware, they have the ambition, but they don't have enough of what their robots actually need to get smarter.
The problem, as Wang Xiaogang explained it, is that building a better robot isn't just about writing better code or designing better joints. Wang, who co-founded SenseTime and now chairs its robotics division Ace Robotics, described the real challenge as something more architectural. You need to connect four things simultaneously—the physical machine itself, the data you collect from it, the artificial intelligence models that learn from that data, and the real-world situations where robots actually have to work. Right now, those four pieces aren't talking to each other the way they need to. They're not locked into a cycle where each one improves the others.
The data problem is particularly acute. Most of the training information robotics companies gather comes from watching humans do tasks—recording their movements, their decisions, their adjustments. But here's where it gets complicated: the way you design a robot's body, the way you collect that human demonstration data, and the physical structure of the machine all need to be optimized together. They can't be developed separately and then bolted together. And even when companies get this right, they're not deploying robots widely enough in actual working environments to generate the volume of real-world data they need. The scenarios where robots operate remain too narrow, too controlled, too limited.
Wang put the central question plainly: how do you take these real-world situations where robots actually learn and scale them up? How do you unlock more scenarios, more variations, more complexity, so that the data keeps flowing?
Yao Maoqing, a senior executive at AgiBot, a Shanghai-based humanoid robot maker, pointed to a different but related constraint. The amount of multi-modal data available about the physical world—the kind of rich, varied information that captures how things actually look, feel, and behave—is nowhere near what the industry has for training large language models. Language models have been fed billions of text samples. Robots trying to understand and navigate physical space don't have anything close to that abundance. This gap is creating a serious bottleneck as companies try to build what they call world models—AI systems that would let next-generation humanoid robots understand their surroundings well enough to move through them, manipulate objects, and adapt to unexpected situations.
The irony is sharp: China's robotics sector has made real progress in mechanical design and hardware engineering. What's holding it back now isn't the metal and motors. It's the information gap. Until companies can figure out how to generate, collect, and share vastly more data about how robots interact with real physical environments—and until they can do it at scale—the robots themselves will keep hitting the same ceiling.
Citas Notables
The key question is how to unlock these scenarios and replicate them at scale— Wang Xiaogang, co-founder of SenseTime and chairman of Ace Robotics
Hardware design for robots must be jointly optimized alongside data-collection methods and physical structure— Wang Xiaogang