Robot Learning from a Physical World Model
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912699138441216 |
|---|---|
| author | Mao, Jiageng He, Sicheng Wu, Hao-Ning You, Yang Sun, Shuyang Wang, Zhicheng Bao, Yanan Chen, Huizhong Guibas, Leonidas Guizilini, Vitor Zhou, Howard Wang, Yue |
| author_facet | Mao, Jiageng He, Sicheng Wu, Hao-Ning You, Yang Sun, Shuyang Wang, Zhicheng Bao, Yanan Chen, Huizhong Guibas, Leonidas Guizilini, Vitor Zhou, Howard Wang, Yue |
| contents | We introduce PhysWorld, a framework that enables robot learning from video generation through physical world modeling. Recent video generation models can synthesize photorealistic visual demonstrations from language commands and images, offering a powerful yet underexplored source of training signals for robotics. However, directly retargeting pixel motions from generated videos to robots neglects physics, often resulting in inaccurate manipulations. PhysWorld addresses this limitation by coupling video generation with physical world reconstruction. Given a single image and a task command, our method generates task-conditioned videos and reconstructs the underlying physical world from the videos, and the generated video motions are grounded into physically accurate actions through object-centric residual reinforcement learning with the physical world model. This synergy transforms implicit visual guidance into physically executable robotic trajectories, eliminating the need for real robot data collection and enabling zero-shot generalizable robotic manipulation. Experiments on diverse real-world tasks demonstrate that PhysWorld substantially improves manipulation accuracy compared to previous approaches. Visit \href{https://pointscoder.github.io/PhysWorld_Web/}{the project webpage} for details. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2511_07416 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Robot Learning from a Physical World Model Mao, Jiageng He, Sicheng Wu, Hao-Ning You, Yang Sun, Shuyang Wang, Zhicheng Bao, Yanan Chen, Huizhong Guibas, Leonidas Guizilini, Vitor Zhou, Howard Wang, Yue Robotics Artificial Intelligence Computer Vision and Pattern Recognition We introduce PhysWorld, a framework that enables robot learning from video generation through physical world modeling. Recent video generation models can synthesize photorealistic visual demonstrations from language commands and images, offering a powerful yet underexplored source of training signals for robotics. However, directly retargeting pixel motions from generated videos to robots neglects physics, often resulting in inaccurate manipulations. PhysWorld addresses this limitation by coupling video generation with physical world reconstruction. Given a single image and a task command, our method generates task-conditioned videos and reconstructs the underlying physical world from the videos, and the generated video motions are grounded into physically accurate actions through object-centric residual reinforcement learning with the physical world model. This synergy transforms implicit visual guidance into physically executable robotic trajectories, eliminating the need for real robot data collection and enabling zero-shot generalizable robotic manipulation. Experiments on diverse real-world tasks demonstrate that PhysWorld substantially improves manipulation accuracy compared to previous approaches. Visit \href{https://pointscoder.github.io/PhysWorld_Web/}{the project webpage} for details. |
| title | Robot Learning from a Physical World Model |
| topic | Robotics Artificial Intelligence Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2511.07416 |