PlayWorld: Learning Robot World Models from Autonomous Play
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866913006543175680 |
|---|---|
| author | Yin, Tenny Mei, Zhiting Zheng, Zhonghe Yamane, Miyu Wang, David Sceats, Jade Bateman, Samuel M. Zha, Lihan Badithela, Apurva Shorinwa, Ola Majumdar, Anirudha |
| author_facet | Yin, Tenny Mei, Zhiting Zheng, Zhonghe Yamane, Miyu Wang, David Sceats, Jade Bateman, Samuel M. Zha, Lihan Badithela, Apurva Shorinwa, Ola Majumdar, Anirudha |
| contents | Action-conditioned video models offer a promising path to building general-purpose robot simulators that can improve directly from data. Yet, despite training on large-scale robot datasets, current state-of-the-art video models still struggle to predict physically consistent robot-object interactions that are crucial in robotic manipulation. To close this gap, we present PlayWorld, a simple, scalable, and fully autonomous pipeline for training high-fidelity video world simulators from interaction experience. In contrast to prior approaches that rely on success-biased human demonstrations, PlayWorld is the first system capable of learning entirely from unsupervised robot self-play, enabling naturally scalable data collection while capturing complex, long-tailed physical interactions essential for modeling realistic object dynamics. Experiments across diverse manipulation tasks show that PlayWorld generates high-quality, physically consistent predictions for contact-rich interactions that are not captured by world models trained on human-collected data. We further demonstrate the versatility of PlayWorld in enabling fine-grained failure prediction and policy evaluation, with up to 40% improvements over human-collected data. Finally, we demonstrate how PlayWorld enables reinforcement learning in the world model, improving policy performance by 65% in success rates when deployed in the real world. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2603_09030 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | PlayWorld: Learning Robot World Models from Autonomous Play Yin, Tenny Mei, Zhiting Zheng, Zhonghe Yamane, Miyu Wang, David Sceats, Jade Bateman, Samuel M. Zha, Lihan Badithela, Apurva Shorinwa, Ola Majumdar, Anirudha Robotics Artificial Intelligence Action-conditioned video models offer a promising path to building general-purpose robot simulators that can improve directly from data. Yet, despite training on large-scale robot datasets, current state-of-the-art video models still struggle to predict physically consistent robot-object interactions that are crucial in robotic manipulation. To close this gap, we present PlayWorld, a simple, scalable, and fully autonomous pipeline for training high-fidelity video world simulators from interaction experience. In contrast to prior approaches that rely on success-biased human demonstrations, PlayWorld is the first system capable of learning entirely from unsupervised robot self-play, enabling naturally scalable data collection while capturing complex, long-tailed physical interactions essential for modeling realistic object dynamics. Experiments across diverse manipulation tasks show that PlayWorld generates high-quality, physically consistent predictions for contact-rich interactions that are not captured by world models trained on human-collected data. We further demonstrate the versatility of PlayWorld in enabling fine-grained failure prediction and policy evaluation, with up to 40% improvements over human-collected data. Finally, we demonstrate how PlayWorld enables reinforcement learning in the world model, improving policy performance by 65% in success rates when deployed in the real world. |
| title | PlayWorld: Learning Robot World Models from Autonomous Play |
| topic | Robotics Artificial Intelligence |
| url | https://arxiv.org/abs/2603.09030 |