SciGroveBeta
Robotics

-0: A Latent Predictive World Action Model for Concurrent Humanoid Loco-Manipulation

Zhe Li, Zhenzhe Zhang, Yangyang Wei, Wenjie Zhang, Xichen Yuan, Peiyuan Zhi, Gen Li, Xinying Guo, Fengjie Gao, Jianfei Yang, Shanghang Zhang

Featured August 21, 2026

AI-generated analysis — This is SciGrove's AI interpretation of the paper, not peer-reviewed content. Always refer to the original paper.

Simply

A new robot system called -0 helps humanoids do complex household tasks like cleaning and moving at the same time by predicting what the world will look like next, using this foresight to generate smooth whole-body actions without needing to draw future videos.

In depth
The paper introduces -0, a novel approach for humanoid robots to perform complex household tasks requiring simultaneous movement and manipulation. It achieves this by learning a latent predictive world-action representation that directly generates controller-compatible whole-body action latents. Instead of reconstructing future videos, -0 uses compact future observation embeddings as a lightweight predictive objective, coupling this visual foresight with diffusion-based whole-body action generation to enable seamless concurrent loco-manipulation.

Key Takeaways

  • 1
    The authors introduce -0, a latent predictive world-action model that couples future visual embedding prediction with diffusion-based action generation for concurrent humanoid loco-manipulation.
  • 2
    A staged training pipeline is developed to learn action-aware visual-language representations and ground human/public visual-motion priors into robot-executable action latents via SONIC-based simulation replay.
  • 3
    The study presents \omega -HOME, a 40-hour real-world multimodal household humanoid dataset, enabling a single -0 model to autonomously execute 11 diverse household tasks with smooth whole-body coordination.

Conceptual Flow

HIGH LEVEL
1
Methodology: How \omega -0 Learns Coordinated Actions

The robot learns to move and manipulate at the same time by predicting what it will see next and using that guess to make smooth, full-body movements.

Language Goal
Current View
Robot State
Predict Future & Plan Actions
Smooth Robot Actions
2
Results: \omega -0 Excels in Real-World Tasks

This new method helps the robot successfully complete many different household tasks, like cleaning and moving objects, much better than older robot systems.

Old Robot Methods
Perform Household Tasks
Lower Success
Higher Success

This breakdown was generated by SciGrove. Get the same analysis — intuition, storyboard, peer review, a runnable prototype and a glossary — on any paper you upload or paste a DOI for.