GigaWorld-0: World Models as Data Engine to Empower Embodied AI
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , , , , , , , , , , , , , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866911293676453888 |
|---|---|
| author | GigaWorld Team Ye, Angen Wang, Boyuan Ni, Chaojun Huang, Guan Zhao, Guosheng Li, Haoyun Zhu, Jiagang Li, Kerui Xu, Mengyuan Deng, Qiuping Wang, Siting Qin, Wenkang Chen, Xinze Wang, Xiaofeng Wang, Yankai Cao, Yu Chang, Yifan Xu, Yuan Ye, Yun Wang, Yang Zhou, Yukun Zhang, Zhengyuan Dong, Zhehao Zhu, Zheng |
| author_facet | GigaWorld Team Ye, Angen Wang, Boyuan Ni, Chaojun Huang, Guan Zhao, Guosheng Li, Haoyun Zhu, Jiagang Li, Kerui Xu, Mengyuan Deng, Qiuping Wang, Siting Qin, Wenkang Chen, Xinze Wang, Xiaofeng Wang, Yankai Cao, Yu Chang, Yifan Xu, Yuan Ye, Yun Wang, Yang Zhou, Yukun Zhang, Zhengyuan Dong, Zhehao Zhu, Zheng |
| contents | World models are emerging as a foundational paradigm for scalable, data-efficient embodied AI. In this work, we present GigaWorld-0, a unified world model framework designed explicitly as a data engine for Vision-Language-Action (VLA) learning. GigaWorld-0 integrates two synergistic components: GigaWorld-0-Video, which leverages large-scale video generation to produce diverse, texture-rich, and temporally coherent embodied sequences under fine-grained control of appearance, camera viewpoint, and action semantics; and GigaWorld-0-3D, which combines 3D generative modeling, 3D Gaussian Splatting reconstruction, physically differentiable system identification, and executable motion planning to ensure geometric consistency and physical realism. Their joint optimization enables the scalable synthesis of embodied interaction data that is visually compelling, spatially coherent, physically plausible, and instruction-aligned. Training at scale is made feasible through our efficient GigaTrain framework, which exploits FP8-precision and sparse attention to drastically reduce memory and compute requirements. We conduct comprehensive evaluations showing that GigaWorld-0 generates high-quality, diverse, and controllable data across multiple dimensions. Critically, VLA model (e.g., GigaBrain-0) trained on GigaWorld-0-generated data achieve strong real-world performance, significantly improving generalization and task success on physical robots without any real-world interaction during training. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2511_19861 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | GigaWorld-0: World Models as Data Engine to Empower Embodied AI GigaWorld Team Ye, Angen Wang, Boyuan Ni, Chaojun Huang, Guan Zhao, Guosheng Li, Haoyun Zhu, Jiagang Li, Kerui Xu, Mengyuan Deng, Qiuping Wang, Siting Qin, Wenkang Chen, Xinze Wang, Xiaofeng Wang, Yankai Cao, Yu Chang, Yifan Xu, Yuan Ye, Yun Wang, Yang Zhou, Yukun Zhang, Zhengyuan Dong, Zhehao Zhu, Zheng Computer Vision and Pattern Recognition Robotics World models are emerging as a foundational paradigm for scalable, data-efficient embodied AI. In this work, we present GigaWorld-0, a unified world model framework designed explicitly as a data engine for Vision-Language-Action (VLA) learning. GigaWorld-0 integrates two synergistic components: GigaWorld-0-Video, which leverages large-scale video generation to produce diverse, texture-rich, and temporally coherent embodied sequences under fine-grained control of appearance, camera viewpoint, and action semantics; and GigaWorld-0-3D, which combines 3D generative modeling, 3D Gaussian Splatting reconstruction, physically differentiable system identification, and executable motion planning to ensure geometric consistency and physical realism. Their joint optimization enables the scalable synthesis of embodied interaction data that is visually compelling, spatially coherent, physically plausible, and instruction-aligned. Training at scale is made feasible through our efficient GigaTrain framework, which exploits FP8-precision and sparse attention to drastically reduce memory and compute requirements. We conduct comprehensive evaluations showing that GigaWorld-0 generates high-quality, diverse, and controllable data across multiple dimensions. Critically, VLA model (e.g., GigaBrain-0) trained on GigaWorld-0-generated data achieve strong real-world performance, significantly improving generalization and task success on physical robots without any real-world interaction during training. |
| title | GigaWorld-0: World Models as Data Engine to Empower Embodied AI |
| topic | Computer Vision and Pattern Recognition Robotics |
| url | https://arxiv.org/abs/2511.19861 |