Omega-0 model reaches 81.8 percent on 11 real home tasks
A humanoid robot walks, looks and works at the same time. That is the job Omega-0 was built to learn, and its creators report 81.8 percent success on 11 real-world home tasks—ahead of pi-0.5, EgoVLA, GR00T-N1.7 and psi-0. Researchers came from Nanyang Technological University, Peking University, HKUST (Guangzhou) and the Beijing Academy of Artificial Intelligence, or BAAI.
The model's central choice is what it predicts. Instead of forecasting future pixels, Omega-0 forecasts future visual features—information about what the robot is likely to see—then feeds those features into a diffusion transformer, which generates coordinated actions for the whole body. The system combines vision, language and proprioception, the robot's own measurements of its body, so an instruction and a changing environment can shape the same action plan.
The training base is Omega-HOME: 40.3 hours of real-environment data, 4,827 trajectories and 24 task types. The demonstrations include first-person and third-person video, language instructions, proprioception and full-body motion. They cover eight household categories, including fetching objects, cleaning surfaces, using appliances and tidying; the trajectories were tele-operated and turned into training signals.
The researchers also designed the evaluation to address a basic concern in robot learning: whether the test tasks have already appeared in training. They removed the 11 fine-tuning tasks from Omega-HOME and combined the remaining data with public human demonstrations for an earlier training stage. That makes the reported comparison more meaningful, while the result remains a research benchmark rather than proof of reliable domestic autonomy.
So what changes in practice? Omega-0 offers a clearer recipe for robots that must continuously sense a home and coordinate their entire body, not just move one arm toward a known target. It could help researchers test household abilities such as fetching, cleaning and appliance use with a common training setup. But long-horizon autonomy, safe interaction, generalizing to new tasks and adapting to complex environments are still open, and the team describes reliable laundry-folding as a distant goal.
Kommentare
Der Thread wird geladen…
Melden Sie sich an, um einen Kommentar zu schreiben. Anmelden