LOADING

Understanding the Galactic Universal, one understands the future of embodied intelligence

How should ChatGPT moments with embodied intelligence be defined?

On August 19th, at the World Robot Conference (WRC), at the “Evolutionary Path of Embodied Intelligent Large Models” sub forum hosted by Galaxy General Motors, the founder and CTO of Galaxy General Motors, Wang He, threw this question to the scene.

 

His answer is that robots can achieve zero sample generalization even when faced with common skills that have not been specifically learned; Ordinary users do not need algorithmic background to teach it new actions. To achieve this, humanoid robots require a brain that can understand the physical world, control their entire body and hands, and continuously learn.

On the newly released small humanoid robot Galbot ET1 (also known as “Galactic Starboy”) by Galaxy General Motors, this concept has a new body. It is reported that “Galaxy Star Boy” is the world’s first intelligent humanoid robot with autonomous learning ability. It is equipped with the AstraBrain Agent developed by Galaxy General Motors, which can observe the environment in real time, understand human instructions, and dynamically plan behavior. Humans do not need to prepare action videos to learn new action skills through interactive teaching.

 

Behind the agile body of the “Galactic Star Boy” is the latest evolution of its embodied large model – the Galactic Star Brain AstraBrain. While the embodied intelligence industry is still fiercely competing around physical performance, Galaxy General Motors has placed heavier stakes on embodied large models. And this model is also determining the upper limit of robots.

This is also the reason why Galaxy General Motors has become the most closely watched company in this year’s WRC.

When the agent enters the physical world
The reason why ‘Galaxy Star Boy’ has sparked discussion is that it carries a new product logic.

In the past two years, Agent has become one of the hottest concepts in the AI industry. Agents in the digital world can call tools between software to help users search for information, and can also perform online tasks for people. Physical intelligent agents face a constantly changing real-world environment. Whether it is changes in the shape of objects or changes in light and position, it is necessary for the model to quickly understand before driving the body to respond.

Galaxy General defines AstraBrain Agent as a ‘physical world native intelligent agent’. Its perception and planning are directly oriented towards the real space, and the result of its thinking ultimately becomes bodily movements. After the robot performs actions, the environment changes accordingly, and new feedback will affect the next step of judgment.

The interactive self-learning ability of “Galaxy Star Boy” is generated under this logic. It can recognize the real-time movement trajectory of human dancers and follow them to complete street dance movements. It can still maintain physical stability in the face of difficult movements such as ground support and handstands. This dynamic physical performance is supported by the AstraBrain BBB, a universal cerebellum in the Milky Way.

This model has over 100000 hours of human action data behind it. One part comes from high-precision motion capture, and the other part is converted from Internet human video. With the help of this data, the model can learn how humans maintain balance, coordinate joints, and complete coherent movements, and then transfer their mobility to robots without the need to write a complete set of motion trajectories in advance.

Behind this is the evolution of embodied large model technology. In the past, adding a skill to robots usually required a professional team to re collect data and complete training, and the skill was also easily bound to specific ontologies and scenarios. AstraBrain Agent hopes to retain the universal capabilities already acquired by the model and adapt to new tasks with minimal interaction. Every time a robot learns a new skill, it continues to expand the same brain.

Around the Galactic Brain, Galactic Universal has constructed a whole brain architecture that includes the brain, pons, and cerebellum. The brain is responsible for understanding the environment and planning actions, the cerebellum is responsible for controlling the whole body and hands, and the brainstem transmits high-level intentions to specific actions.

Among them, the brain adopts the world action model WAM architecture first proposed by Galaxy General. VLA can generate actions based on visual and linguistic cues, while world models excel at deducing how the environment will change. WAM brings the two capabilities into the unified model, so that the robot can understand the possible physical results of an action and decide the next action accordingly.

The core breakthrough of AstraBrain WAM is to integrate cross ontology, cross scenario, and multi tasking capabilities into the same set of models. The same brain can drive the G1 robot to use dexterous hands to pick up goods from supermarkets, as well as complete flexible operations such as folding clothes; After replacing the robot body, it can also perform brand new container handling tasks. The accumulation of models has freed itself from the limitations of hardware and scenarios, and this extensive ability transfer is the first of its kind in the industry.

The “Galaxy Star Boy” is also equipped with the AstraBrain WCC Universal Cerebellar Basic Module, which was independently developed by Galaxy General Motors. It is trained on large-scale human motion data and can convert the intentions generated by the AstraBrain Agent into stable and coherent body movements. The model also has excellent generalization ability when facing actions that have not appeared during training.

Tennis is one of the most challenging scenarios for testing this set of technical abilities, as this sport has a strong competitive nature and leaves robots with extremely short reaction times. It not only needs to predict the trajectory of the incoming ball, but also adjust its position in a timely manner, and accurately hit the ball while maintaining balance. The “Galaxy Star Boy” is equipped with the Galaxy Universal LateNT High Dynamic Tennis Match Control Framework, which can adjust its movements in real time based on the ball received from a real person in actual demonstrations, completing personified matches.

From following humans to making independent judgments, the industry has gradually formed a new consensus: the upper limit that a humanoid robot can reach increasingly depends on how many new abilities its “brain” can grow behind it. The more bodies it enters, the richer its experience accumulated in the physical world, and the next robot will have a higher starting point as a result.

A real test of a brain
At the Galaxy Universal booth, this’ Galaxy Star Brain ‘was also put into different tasks for testing.

The most on-site experience is a seemingly ordinary breakfast task: the robot needs to continuously complete steps such as picking up bread, pouring water, and arranging dishes. During the execution process, there may be occasional situations, such as the cup being taken away by the audience or the target object suddenly being obstructed. But in the on-site demonstration, the robot will re plan according to the new state in front of it and complete the remaining work with ease.

Traditional robots execute preset processes in a fixed environment, and as long as the conditions of a certain step change, the entire task is easily interrupted. However, in the real world, strict adherence to pre written scripts is rare. A person reaching out at the table and a cup being temporarily moved away will change what the robot should do next second.

The difficulty of long-range tasks lies in these temporary changes. The longer the action lasts, the more likely the errors generated earlier will affect the subsequent steps. Robots need to constantly update their understanding of the environment while maintaining memory of the entire task during execution.

Another high difficulty scenario is folding clothes, which tests the embodied large model’s understanding of flexible objects. Clothing does not have a fixed shape, and the folds of the fabric change every time the robot grabs it. The model needs to re evaluate the current state of the clothes and find suitable grasping points. Whether the task can continue depends on whether the model can update its judgment after each operation step.

Of course, for real industrial scenarios, a successful demonstration is far from enough. The accuracy of repetitive tasks every day determines whether customers are willing to continue paying for robots.

Galaxy General Motors is the most qualified embodiment manufacturer to answer this question. This company has already applied its products on a large scale to smart pharmacies and instant retail. Robots need to accurately find targets among tens of thousands of products, complete grabbing, and hand them over to riders. This type of task is repeated a large number of times every day, and its accuracy, stability, and ability to handle exceptions directly affect the business.

 

In industrial scenarios, the Galabot S1 from Galaxy Universal has a maximum load capacity of 50 kilograms on both arms, which can handle heavy material handling. After the load increases, the robot needs to readjust its body posture and output mode, and also determine the position of surrounding personnel to ensure collaborative safety.

The robot forms used in different scenarios are different, and the problems faced also vary greatly. On one side are soft and easily deformed clothes, on the other side are dense shelves, and on the other side are industrial heavy objects. They share the technological foundation of the Galactic Brain and also verify its generalization ability. And this is precisely the starting point for the universality of embodied models.

For Galaxy General, there is another layer of value for models entering the industry. Smart pharmacies, instant retail, and industrial production lines will continue to generate real data. These data record issues that are difficult to cover in the simulation environment, such as changes in packaging materials, deviations caused by long-term equipment operation, and random interference from on-site personnel. After a new round of training, the updated abilities of the model will return to the robot.

At the forum, Galaxy General Motors also announced that it will open simulation platforms, data collection devices, embodied models, and reinforcement learning training pipelines to technology companies and industry partners. Ordinary developers can also create new actions and applications around the “Galaxy Star Boy”. This is also the necessary path for embodied models to move towards universality: as the ecosystem becomes more open and the number of participants increases, the real problems that big models come into contact with become more diverse, and the speed of model evolution will also accelerate accordingly.

Only when models enter the real world can technological achievements be transformed into sustainable productivity. Galaxy General has demonstrated in this year’s WRC that the Galactic Brain already possesses such capability.

The embodied large model is the new watershed
In the past few years, the spotlight of the humanoid robot industry has mostly been on the body.

Yushu Technology, which has demonstrated outstanding performance in robot operation control and engineering capabilities, made its debut on August 19th as an example. On the opening day, Yushu Technology achieved a growth of over 629%, with a market value of 444.9 billion yuan at one point. The capital market has offered a high price for the commercial value of humanoid robots with an extremely enthusiastic debut.

Robots running fast, jumping high, and withstanding huge impacts are indeed the most intuitive scales of technological progress. But when robots enter pharmacies, supermarkets, and factories, an “active” robot is obviously no longer able to meet the demand. It needs to understand a vague instruction, handle objects that have not been seen during training, and deal with personnel movement and object displacement. The body determines where a robot can reach, while the brain determines its ability boundaries.

The embodied large model has thus become a new watershed in the industry.

The performance of the ontology can be accelerated through supply chain and engineering investment, while the growth cycle of a universal brain is often longer. It needs to learn physical laws in a large number of tasks and enter real scenarios for verification. As the tasks experienced by the model become more diverse, the experience that can be called upon when dealing with new problems will also increase. Even if later generations use similar hardware, it is difficult to quickly supplement this know-how.

The listing of Yushu demonstrates the commercial value of robot bodies, while the competition for embodied models is defining the next stage of the industry’s capability coordinates.

Galaxy General has long regarded data infrastructure as the core of embodied big models. For this purpose, a five layer data pyramid was constructed based on the number of stars in the Milky Way. Among them, Internet data helps model understand semantics, human action data provides operating experience, simulation platform can generate training samples on a large scale, and real machine teleoperation data is responsible for calibrating fine actions. After the robot enters the actual scene, the problems generated during work will return to the training system.

Wang He disclosed on the forum that Galaxy General Motors has accumulated 1 million hours of human data and 80000 hours of real-world scenario data. As early as 2021, the team began building a first person perspective human object interaction dataset. Nowadays, data collection devices, simulation platforms, and model evaluation systems have been integrated into the same infrastructure.

When more robots enter real scenes, the Galactic Brain can gain richer physical experience; After the model capability is improved, the adaptation cost of new robots and tasks will also decrease. The deployment scale and model capability mutually drive each other, forming a constantly growing data flywheel.

Returning to the most discussed question of this year’s WRC: When will the ChatGPT moment of embodied intelligence arrive?

The standard given by Wang He is that robots can achieve zero sample generalization when facing common skills that humans do not need to learn specifically, with a success rate of 70% to 80%; Ordinary users can also complete post training at low cost, allowing the model to quickly adapt to their work environment. At this point, embodied intelligence truly has a popular foundation similar to ChatGPT.

 

Reaching 70-80% model capability is just the first step. In reality, customers still require algorithm engineers to collect robot motion data, and then complete annotation and debugging. The high cost of adaptation will keep robots out of the door of a large number of small and medium-sized enterprises and ordinary users. To popularize embodied models, we also need to solve the problem of ‘anyone can teach and use’.

Wang He also introduced the training plan for WAM-TTT testing in his speech at this forum. Users wear a first person perspective camera to capture their own work process, and the model can be deployed using unlabeled videos without the need to re collect robot motion data.

This further advances the generalization of embodied intelligence. The unified model is responsible for accumulating general knowledge of the physical world, and humans only need to demonstrate their work, and robots have the opportunity to quickly acquire corresponding abilities. When the process of teaching robots is simple enough, the embodied large model can truly enter various industries.

At the end of the forum speech, Wang He outlined a broader industrial landscape: future robots may have a shipment scale close to that of mobile phones, maintain automotive level product value, and inherit the ability to continuously upgrade large models. A set of models completes evolution, and countless robots can gain ability enhancement. The scale benefits brought by software will open up a trillion dollar market.

Humans have spent decades accumulating language and images for the digital world, while the training corpus for the physical world still needs to be personally written by robots. They will identify tens of thousands of products in pharmacies, handle accidents in factories, and learn a new job through demonstrations by ordinary people. Every real action becomes the starting point of the next evolution.

When the same brain can enter different bodies, and teaching robots a skill becomes as natural as teaching humans, the ChatGPT moment of embodied intelligence will truly arrive. A brain that can share experiences and continuously evolve among thousands of bodies will allow AI to truly enter the physical world and open the gateway to embodied AGI.

© 版权声明

相关文章