WRC2026 observation: “Bright line” represents ground, “dark line” represents data
At this year’s WRC, humanoid robots were sorting packages on conveyor belts and screwing screws at workstations. This kind of “doing practical things” scene, like the still lively performance exhibitions, has become the norm, which is a gratifying industry progress for many people.
But looking back, this is not the entirety of WRC 2026, at least only a part of the ‘clear line’.
The amount of data, especially the amount of high-quality data, corresponds to the amount of AI capability. ”The statement made by Wang Xingxing, founder of Yushu Technology, at WRC is almost an unspoken consensus among all manufacturers at this conference. The limbs of the humanoid robot have been trained to run, jump, somersault, and climb stairs, but the “brain” is still hungry.
The question is, where does the data fed to the brain come from?
A set of data repeatedly cited on WRC shows that the current compliance data for real physical interaction scenarios in China is only 500000 hours, while the commercialization of robots requires tens of millions of hours, with a gap of over 99%.
From 100000 to millions, there is a difference of two orders of magnitude in between.
So, the entire WRC, ‘data’ became a dark line.
Players from all walks of life demonstrate their abilities, some open source human behavior data, some sell data collection gloves, some bet on simulation, and some insist on real machine data.
In this hustle and bustle, there are some embodied intelligence companies that have no booths and are hardly noticed, but hold the rarest batch of data in the entire industry.
Open source human data, gamble on the mass line
Lightwheel Intelligence is one of the biggest players in terms of data level actions at this year’s WRC.
On August 20th, they released EgoSuite Open100K, the world’s first open-source dataset of 100000 hour full modal human behavior.
The scale of this dataset is 100000 hours of first person human data, covering 7 categories of environments, 128 scenarios, over 15000 collection scenarios, and more than 15000 tasks.
The data is mainly based on the first perspective of the headset, with some wrist cameras added, providing hand posture, full body posture, and event level semantic annotation. The first batch of data has been opened for download through Hugging Face.
Yang Haibo, CEO of Lightwheel Intelligence, explained why it is necessary to promote such a data route. He believes that relying solely on real machine remote operation is difficult to support the supply of training materials at the level of millions of hours or more. The industry urgently needs two parallel paths: human video data and simulation synthesis data.
So the strategy of Lightwheel is the “mass line” – open source human behavior data, build a Real2Sim2Real continuous learning loop, and even propose a 5-year 10 billion hour embodied intelligent data co construction plan – in plain terms, if I can’t finish it alone, let the whole industry take it together.
Real data+simulation parallel, walking on two legs
At present, the combination of a small proportion of real data and a large proportion of simulation data is a temporary solution in the absence of data, and has also become the main focus of many industrial chain enterprises.
Xinghai Tu is a representative of this faction. CEO Gao Jiyang made it clear at WRC that he insists on prioritizing the use of real data. After calculating the cost of data, computing power, and R&D engineer labor, he concluded that “the cost of data is actually relatively controllable and not as high as imagined. The most expensive thing is actually the time of R&D engineers
The real data base of Xinghai Map includes the open source GOD real scene dataset from September 2025, which has the highest download volume in the world; Expand data sources by investing in companies such as Jianzhi and Yuanliu; We plan to achieve a real data scale of 1 million hours by 2026.
However, at the same time, Xinghai Map has not given up on simulation, and its paper citation rate ranks first in the world model field.
In a sense, Gao Jiyang’s statement is more like an ideal. Currently, simulation data is still the mainstay of embodied intelligence training. In the future, everyone wants to use as much real data as possible, but currently there is no such condition.
Even if the popular Crispy fried chicken Yushu Technology makes the robot further adapt to the physical world through the real robot operation data, the massive Internet data for pre training is unavoidable.
Walking on two legs, one responsible for mind filling and broadening one’s horizons, and the other responsible for being down-to-earth, will be the norm for the long-term development of embodied intelligence.
Run out of the mine with real data on extreme working conditions
By now, you may have discovered an issue where almost all player data, whether it is human behavior videos, simulation synthesis, or real machine remote operations, comes from relatively standardized, structured, safe and controllable scenarios, such as laboratories, factory workshops, home environments, and so on.
But there is a type of data that none of the above methods can collect.
That is the real physical interaction data under extreme working conditions – in a mine 1000 meters underground, next to a smelting furnace at thousands of degrees Celsius, at the edge of the blasting zone in an open-pit mine, and at a strong mixing and weaving operation site under dust and light interference.
These scenarios, which humans are unwilling to go to (63% of mining accidents are caused by human factors), cannot be simulated (the physical complexity of unstructured extreme working conditions exceeds the modeling ability of simulation engines), and cannot be remotely operated (communication delay, occlusion attenuation, multi vehicle channel congestion).
What Xidi Zhijia is holding is precisely this batch of data.
In the keynote speech at WRC, CEO of Xidi Intelligent Driving, Husobo, presented a set of numbers: as of the first half of 2026, Xidi’s autonomous driving fleet has exceeded 3400 units, of which more than 1900 unmanned mining trucks have been shipped, covering nearly 40 mines worldwide, with a maximum normal operation scale of 220 units per mine.
What does this mean? This means that 1900 heavy machines operate in over 40 extreme working conditions in mines every day, continuously generating first-hand data for perception, decision-making, and execution throughout the entire chain. This is not clean data from the laboratory, not synthetic data from the simulator, not behavior videos taken by humans wearing headbands – this is data produced by machines in real hellish environments.
Hu Sibo said a crucial sentence in his speech, “Autonomous driving not only accumulates algorithms and mileage, but also the cognition of 1900 cars running out of more than 40 mines every day, and more importantly, the understanding of the physical world by machines
The implicit meaning of this sentence is that the overloaded embodied intelligence moat is not a model parameter, but a real data accumulation under extreme working conditions.
Based on this batch of data, Xidi has built a three-layer architecture heavy-duty embodied intelligent technology platform:
The bottom layer is a unified data engine that aggregates multidimensional data from production, environment, and vehicles;
In the middle is the overloaded world model, the terminal action model is responsible for millisecond level real-time decision-making, and the cloud training model continues to iterate;
The upper layer is the cluster decision-making and scheduling system, which can schedule over 1000 intelligent agents in a single scenario.
More importantly, Xidi has not stopped at the “transportation” stage. 70% of mining operations are non transportation – drilling, blasting, excavation, and loading. Xidi is extending the “brain” trained from transportation data to the entire process of operations: 34 mining robots have been shipped, open-pit mining loading robots have achieved unmanned drilling, loading, and vehicle moving in all stages, smelting tank holding robots transport 75 tons of slag tanks at high temperatures of thousands of degrees Celsius, and underground mining excavation robots have completed complete machine manufacturing.
The business fundamentals that support all of this are a set of data that appears particularly “dazzling” in the field of embodied intelligence: the total revenue in the first half of 2026 was 804 million yuan, a year-on-year increase of 97%; Gross profit of 212 million yuan, a year-on-year increase of 204%; The revenue from autonomous driving increased by 107.3% year-on-year.
Against the backdrop of high investment and burning money in the entire industry, Xidi is one of the few players who has already run through the commercial closed loop and has a clear profit path.
Xidi’s positioning is also very clear, not as a terminal, only as a brain, and as a technology partner for intelligent heavy-duty machines.
This “brain supplier” model allows it to take on lightweight tasks, achieving the ultimate in algorithms, data, and cluster scheduling, while expanding along three lines: carriers, scenarios, and going global. It collaborates with Proton Automotive and Cloud Deep Technology, extending from mines to smelting and logistics parks, and bringing Chinese solutions to resource-based countries such as Australia, Brazil, and Indonesia.
Data mining hardware, selling mining shovels is also a good business
If data is gold, then this group sells mining shovels.
At the WRC site, data acquisition hardware manufacturers collectively erupted.
Octopus Power has launched three products: a fisheye headband, an electromyographic wristband, and an exoskeleton isomorphic data collection glove, collectively known as OctoSense. The core selling point of OctoSense is the world’s first achievement of zero sample generalization of electromyography across individuals. The background is that the traditional difficulty of electromyography collection is the large difference in electromyographic signals between individuals. Octopus Power attempts to align human operation data with real machines almost without loss.
The independent variable also set up a demonstration area for data collection without ontology on site, with three data collection hardware unified data output standards, and the collected materials can be seamlessly integrated into their own data service pipeline.
DaXiao Robotics brings the Environmental Data Collection Solution 2.0, which uses ACE Ego Kit, ACE Data Engine, and ACE Ego Matrix to build a complete chain from collection, automatic annotation to cross ontology applications.
BrainCo’s approach is more comprehensive, integrating three data sources: real machine execution, human teaching, and simulation generation. The dual arm wheeled data acquisition platform synchronously collects visual, tactile, and motion states, while the exoskeleton human data acquisition glove records the human operation process, transforming human motion experience into data for robot learning.
Of course, there are also some more cutting-edge approaches.
For example, Fu Liye proposed the “brain computer data acquisition” model, which establishes a data system around the brain, human, and machine. The core purpose is to compare the differences in EEG under the three states of “execution, imagination, and teleoperation robots”, providing a foundation for research on motor imagination. QiangBrain Technology directly demonstrated on site how brain computer interfaces connect humans and robots: humans generate action intentions, sensors collect EEG signals related to tasks, convert them into instructions that can be executed by the robot control system, and drive humanoid robots to complete corresponding actions.
Kankan Intelligence has also incorporated trendy smart glasses and launched an ultra lightweight car grade precision heterogeneous data acquisition glasses weighing only 56 grams, focusing on passive massive data acquisition in real scenarios.
Regardless of the approach, the characteristic of this faction is to try to build the infrastructure for data collection. Its logic is simple: if you all lack data, then I will sell you the tools for data collection.
Conclusion
At WRC Developer Night, someone proposed a five layer data progression framework: human demonstrations provide knowledge priors, real machine execution provides real-world experience, simulation provides exploration efficiency, failed data provides boundary breakthroughs, and industrial scenarios provide sustainable closed loops.
Among these five layers, the rarest and most valuable are the last two layers – failure data and industry scenario closed-loop data. Because the first three layers can be scaled up through open source, hardware, and simulation, while the last two layers can only be run on the industrial site with real tools and guns.
From this perspective, the data battle of WRC 2026 has actually been divided into levels. Lightwheels are competing for the breadth of data scale, octopuses are competing for the accuracy of data acquisition tools, Xinghai Maps are competing for the balance between reality and simulation, while Xidi Intelligent Driving is competing for a track that others cannot enter at all, and industrial closed-loop data under extreme working conditions.
While more than 300 companies are teaching robots how to grab, sort, and screw in the exhibition hall of Yizhuang, 1900 unmanned mining trucks are continuously producing the rarest batch of “dirty data” in the entire industry in 40 mines, in an environment of dust, high temperature, and strong mixing.
The ultimate goal of humanoid robots is to enter households, but before that, a group of machines must first enter hell. And the data that runs out of hell is the real moat that cannot be replicated.