Robot Learning from a Physical World Model

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mao, Jiageng, He, Sicheng, Wu, Hao-Ning, You, Yang, Sun, Shuyang, Wang, Zhicheng, Bao, Yanan, Chen, Huizhong, Guibas, Leonidas, Guizilini, Vitor, Zhou, Howard, Wang, Yue
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912699138441216
author Mao, Jiageng
He, Sicheng
Wu, Hao-Ning
You, Yang
Sun, Shuyang
Wang, Zhicheng
Bao, Yanan
Chen, Huizhong
Guibas, Leonidas
Guizilini, Vitor
Zhou, Howard
Wang, Yue
author_facet Mao, Jiageng
He, Sicheng
Wu, Hao-Ning
You, Yang
Sun, Shuyang
Wang, Zhicheng
Bao, Yanan
Chen, Huizhong
Guibas, Leonidas
Guizilini, Vitor
Zhou, Howard
Wang, Yue
contents We introduce PhysWorld, a framework that enables robot learning from video generation through physical world modeling. Recent video generation models can synthesize photorealistic visual demonstrations from language commands and images, offering a powerful yet underexplored source of training signals for robotics. However, directly retargeting pixel motions from generated videos to robots neglects physics, often resulting in inaccurate manipulations. PhysWorld addresses this limitation by coupling video generation with physical world reconstruction. Given a single image and a task command, our method generates task-conditioned videos and reconstructs the underlying physical world from the videos, and the generated video motions are grounded into physically accurate actions through object-centric residual reinforcement learning with the physical world model. This synergy transforms implicit visual guidance into physically executable robotic trajectories, eliminating the need for real robot data collection and enabling zero-shot generalizable robotic manipulation. Experiments on diverse real-world tasks demonstrate that PhysWorld substantially improves manipulation accuracy compared to previous approaches. Visit \href{https://pointscoder.github.io/PhysWorld_Web/}{the project webpage} for details.
format Preprint
id arxiv_https___arxiv_org_abs_2511_07416
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Robot Learning from a Physical World Model
Mao, Jiageng
He, Sicheng
Wu, Hao-Ning
You, Yang
Sun, Shuyang
Wang, Zhicheng
Bao, Yanan
Chen, Huizhong
Guibas, Leonidas
Guizilini, Vitor
Zhou, Howard
Wang, Yue
Robotics
Artificial Intelligence
Computer Vision and Pattern Recognition
We introduce PhysWorld, a framework that enables robot learning from video generation through physical world modeling. Recent video generation models can synthesize photorealistic visual demonstrations from language commands and images, offering a powerful yet underexplored source of training signals for robotics. However, directly retargeting pixel motions from generated videos to robots neglects physics, often resulting in inaccurate manipulations. PhysWorld addresses this limitation by coupling video generation with physical world reconstruction. Given a single image and a task command, our method generates task-conditioned videos and reconstructs the underlying physical world from the videos, and the generated video motions are grounded into physically accurate actions through object-centric residual reinforcement learning with the physical world model. This synergy transforms implicit visual guidance into physically executable robotic trajectories, eliminating the need for real robot data collection and enabling zero-shot generalizable robotic manipulation. Experiments on diverse real-world tasks demonstrate that PhysWorld substantially improves manipulation accuracy compared to previous approaches. Visit \href{https://pointscoder.github.io/PhysWorld_Web/}{the project webpage} for details.
title Robot Learning from a Physical World Model
topic Robotics
Artificial Intelligence
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.07416