WoW: Towards a World omniscient World model Through Embodied Interaction

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chi, Xiaowei, Jia, Peidong, Fan, Chun-Kai, Ju, Xiaozhu, Mi, Weishi, Zhang, Kevin, Qin, Zhiyuan, Tian, Wanxin, Ge, Kuangzhi, Li, Hao, Qian, Zezhong, Chen, Anthony, Zhou, Qiang, Jia, Yueru, Liu, Jiaming, Dai, Yong, Wuwu, Qingpo, Bai, Chengyu, Wang, Yu-Kai, Li, Ying, Chen, Lizhang, Bao, Yong, Jiang, Zhiyuan, Zhu, Jiacheng, Tang, Kai, An, Ruichuan, Luo, Yulin, Feng, Qiuxuan, Zhou, Siyuan, Chan, Chi-min, Hou, Chengkai, Xue, Wei, Han, Sirui, Guo, Yike, Zhang, Shanghang, Tang, Jian
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909849271402496
author Chi, Xiaowei
Jia, Peidong
Fan, Chun-Kai
Ju, Xiaozhu
Mi, Weishi
Zhang, Kevin
Qin, Zhiyuan
Tian, Wanxin
Ge, Kuangzhi
Li, Hao
Qian, Zezhong
Chen, Anthony
Zhou, Qiang
Jia, Yueru
Liu, Jiaming
Dai, Yong
Wuwu, Qingpo
Bai, Chengyu
Wang, Yu-Kai
Li, Ying
Chen, Lizhang
Bao, Yong
Jiang, Zhiyuan
Zhu, Jiacheng
Tang, Kai
An, Ruichuan
Luo, Yulin
Feng, Qiuxuan
Zhou, Siyuan
Chan, Chi-min
Hou, Chengkai
Xue, Wei
Han, Sirui
Guo, Yike
Zhang, Shanghang
Tang, Jian
author_facet Chi, Xiaowei
Jia, Peidong
Fan, Chun-Kai
Ju, Xiaozhu
Mi, Weishi
Zhang, Kevin
Qin, Zhiyuan
Tian, Wanxin
Ge, Kuangzhi
Li, Hao
Qian, Zezhong
Chen, Anthony
Zhou, Qiang
Jia, Yueru
Liu, Jiaming
Dai, Yong
Wuwu, Qingpo
Bai, Chengyu
Wang, Yu-Kai
Li, Ying
Chen, Lizhang
Bao, Yong
Jiang, Zhiyuan
Zhu, Jiacheng
Tang, Kai
An, Ruichuan
Luo, Yulin
Feng, Qiuxuan
Zhou, Siyuan
Chan, Chi-min
Hou, Chengkai
Xue, Wei
Han, Sirui
Guo, Yike
Zhang, Shanghang
Tang, Jian
contents Humans develop an understanding of intuitive physics through active interaction with the world. This approach is in stark contrast to current video models, such as Sora, which rely on passive observation and therefore struggle with grasping physical causality. This observation leads to our central hypothesis: authentic physical intuition of the world model must be grounded in extensive, causally rich interactions with the real world. To test this hypothesis, we present WoW, a 14-billion-parameter generative world model trained on 2 million robot interaction trajectories. Our findings reveal that the model's understanding of physics is a probabilistic distribution of plausible outcomes, leading to stochastic instabilities and physical hallucinations. Furthermore, we demonstrate that this emergent capability can be actively constrained toward physical realism by SOPHIA, where vision-language model agents evaluate the DiT-generated output and guide its refinement by iteratively evolving the language instructions. In addition, a co-trained Inverse Dynamics Model translates these refined plans into executable robotic actions, thus closing the imagination-to-action loop. We establish WoWBench, a new benchmark focused on physical consistency and causal reasoning in video, where WoW achieves state-of-the-art performance in both human and autonomous evaluation, demonstrating strong ability in physical causality, collision dynamics, and object permanence. Our work provides systematic evidence that large-scale, real-world interaction is a cornerstone for developing physical intuition in AI. Models, data, and benchmarks will be open-sourced.
format Preprint
id arxiv_https___arxiv_org_abs_2509_22642
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle WoW: Towards a World omniscient World model Through Embodied Interaction
Chi, Xiaowei
Jia, Peidong
Fan, Chun-Kai
Ju, Xiaozhu
Mi, Weishi
Zhang, Kevin
Qin, Zhiyuan
Tian, Wanxin
Ge, Kuangzhi
Li, Hao
Qian, Zezhong
Chen, Anthony
Zhou, Qiang
Jia, Yueru
Liu, Jiaming
Dai, Yong
Wuwu, Qingpo
Bai, Chengyu
Wang, Yu-Kai
Li, Ying
Chen, Lizhang
Bao, Yong
Jiang, Zhiyuan
Zhu, Jiacheng
Tang, Kai
An, Ruichuan
Luo, Yulin
Feng, Qiuxuan
Zhou, Siyuan
Chan, Chi-min
Hou, Chengkai
Xue, Wei
Han, Sirui
Guo, Yike
Zhang, Shanghang
Tang, Jian
Robotics
Computer Vision and Pattern Recognition
Multimedia
Humans develop an understanding of intuitive physics through active interaction with the world. This approach is in stark contrast to current video models, such as Sora, which rely on passive observation and therefore struggle with grasping physical causality. This observation leads to our central hypothesis: authentic physical intuition of the world model must be grounded in extensive, causally rich interactions with the real world. To test this hypothesis, we present WoW, a 14-billion-parameter generative world model trained on 2 million robot interaction trajectories. Our findings reveal that the model's understanding of physics is a probabilistic distribution of plausible outcomes, leading to stochastic instabilities and physical hallucinations. Furthermore, we demonstrate that this emergent capability can be actively constrained toward physical realism by SOPHIA, where vision-language model agents evaluate the DiT-generated output and guide its refinement by iteratively evolving the language instructions. In addition, a co-trained Inverse Dynamics Model translates these refined plans into executable robotic actions, thus closing the imagination-to-action loop. We establish WoWBench, a new benchmark focused on physical consistency and causal reasoning in video, where WoW achieves state-of-the-art performance in both human and autonomous evaluation, demonstrating strong ability in physical causality, collision dynamics, and object permanence. Our work provides systematic evidence that large-scale, real-world interaction is a cornerstone for developing physical intuition in AI. Models, data, and benchmarks will be open-sourced.
title WoW: Towards a World omniscient World model Through Embodied Interaction
topic Robotics
Computer Vision and Pattern Recognition
Multimedia
url https://arxiv.org/abs/2509.22642