WoW: Towards a World omniscient World model Through Embodied Interaction
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866909849271402496 |
|---|---|
| author | Chi, Xiaowei Jia, Peidong Fan, Chun-Kai Ju, Xiaozhu Mi, Weishi Zhang, Kevin Qin, Zhiyuan Tian, Wanxin Ge, Kuangzhi Li, Hao Qian, Zezhong Chen, Anthony Zhou, Qiang Jia, Yueru Liu, Jiaming Dai, Yong Wuwu, Qingpo Bai, Chengyu Wang, Yu-Kai Li, Ying Chen, Lizhang Bao, Yong Jiang, Zhiyuan Zhu, Jiacheng Tang, Kai An, Ruichuan Luo, Yulin Feng, Qiuxuan Zhou, Siyuan Chan, Chi-min Hou, Chengkai Xue, Wei Han, Sirui Guo, Yike Zhang, Shanghang Tang, Jian |
| author_facet | Chi, Xiaowei Jia, Peidong Fan, Chun-Kai Ju, Xiaozhu Mi, Weishi Zhang, Kevin Qin, Zhiyuan Tian, Wanxin Ge, Kuangzhi Li, Hao Qian, Zezhong Chen, Anthony Zhou, Qiang Jia, Yueru Liu, Jiaming Dai, Yong Wuwu, Qingpo Bai, Chengyu Wang, Yu-Kai Li, Ying Chen, Lizhang Bao, Yong Jiang, Zhiyuan Zhu, Jiacheng Tang, Kai An, Ruichuan Luo, Yulin Feng, Qiuxuan Zhou, Siyuan Chan, Chi-min Hou, Chengkai Xue, Wei Han, Sirui Guo, Yike Zhang, Shanghang Tang, Jian |
| contents | Humans develop an understanding of intuitive physics through active interaction with the world. This approach is in stark contrast to current video models, such as Sora, which rely on passive observation and therefore struggle with grasping physical causality. This observation leads to our central hypothesis: authentic physical intuition of the world model must be grounded in extensive, causally rich interactions with the real world. To test this hypothesis, we present WoW, a 14-billion-parameter generative world model trained on 2 million robot interaction trajectories. Our findings reveal that the model's understanding of physics is a probabilistic distribution of plausible outcomes, leading to stochastic instabilities and physical hallucinations. Furthermore, we demonstrate that this emergent capability can be actively constrained toward physical realism by SOPHIA, where vision-language model agents evaluate the DiT-generated output and guide its refinement by iteratively evolving the language instructions. In addition, a co-trained Inverse Dynamics Model translates these refined plans into executable robotic actions, thus closing the imagination-to-action loop. We establish WoWBench, a new benchmark focused on physical consistency and causal reasoning in video, where WoW achieves state-of-the-art performance in both human and autonomous evaluation, demonstrating strong ability in physical causality, collision dynamics, and object permanence. Our work provides systematic evidence that large-scale, real-world interaction is a cornerstone for developing physical intuition in AI. Models, data, and benchmarks will be open-sourced. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2509_22642 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | WoW: Towards a World omniscient World model Through Embodied Interaction Chi, Xiaowei Jia, Peidong Fan, Chun-Kai Ju, Xiaozhu Mi, Weishi Zhang, Kevin Qin, Zhiyuan Tian, Wanxin Ge, Kuangzhi Li, Hao Qian, Zezhong Chen, Anthony Zhou, Qiang Jia, Yueru Liu, Jiaming Dai, Yong Wuwu, Qingpo Bai, Chengyu Wang, Yu-Kai Li, Ying Chen, Lizhang Bao, Yong Jiang, Zhiyuan Zhu, Jiacheng Tang, Kai An, Ruichuan Luo, Yulin Feng, Qiuxuan Zhou, Siyuan Chan, Chi-min Hou, Chengkai Xue, Wei Han, Sirui Guo, Yike Zhang, Shanghang Tang, Jian Robotics Computer Vision and Pattern Recognition Multimedia Humans develop an understanding of intuitive physics through active interaction with the world. This approach is in stark contrast to current video models, such as Sora, which rely on passive observation and therefore struggle with grasping physical causality. This observation leads to our central hypothesis: authentic physical intuition of the world model must be grounded in extensive, causally rich interactions with the real world. To test this hypothesis, we present WoW, a 14-billion-parameter generative world model trained on 2 million robot interaction trajectories. Our findings reveal that the model's understanding of physics is a probabilistic distribution of plausible outcomes, leading to stochastic instabilities and physical hallucinations. Furthermore, we demonstrate that this emergent capability can be actively constrained toward physical realism by SOPHIA, where vision-language model agents evaluate the DiT-generated output and guide its refinement by iteratively evolving the language instructions. In addition, a co-trained Inverse Dynamics Model translates these refined plans into executable robotic actions, thus closing the imagination-to-action loop. We establish WoWBench, a new benchmark focused on physical consistency and causal reasoning in video, where WoW achieves state-of-the-art performance in both human and autonomous evaluation, demonstrating strong ability in physical causality, collision dynamics, and object permanence. Our work provides systematic evidence that large-scale, real-world interaction is a cornerstone for developing physical intuition in AI. Models, data, and benchmarks will be open-sourced. |
| title | WoW: Towards a World omniscient World model Through Embodied Interaction |
| topic | Robotics Computer Vision and Pattern Recognition Multimedia |
| url | https://arxiv.org/abs/2509.22642 |