GigaWorld-0: World Models as Data Engine to Empower Embodied AI

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: GigaWorld Team, Ye, Angen, Wang, Boyuan, Ni, Chaojun, Huang, Guan, Zhao, Guosheng, Li, Haoyun, Zhu, Jiagang, Li, Kerui, Xu, Mengyuan, Deng, Qiuping, Wang, Siting, Qin, Wenkang, Chen, Xinze, Wang, Xiaofeng, Wang, Yankai, Cao, Yu, Chang, Yifan, Xu, Yuan, Ye, Yun, Wang, Yang, Zhou, Yukun, Zhang, Zhengyuan, Dong, Zhehao, Zhu, Zheng
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911293676453888
author GigaWorld Team
Ye, Angen
Wang, Boyuan
Ni, Chaojun
Huang, Guan
Zhao, Guosheng
Li, Haoyun
Zhu, Jiagang
Li, Kerui
Xu, Mengyuan
Deng, Qiuping
Wang, Siting
Qin, Wenkang
Chen, Xinze
Wang, Xiaofeng
Wang, Yankai
Cao, Yu
Chang, Yifan
Xu, Yuan
Ye, Yun
Wang, Yang
Zhou, Yukun
Zhang, Zhengyuan
Dong, Zhehao
Zhu, Zheng
author_facet GigaWorld Team
Ye, Angen
Wang, Boyuan
Ni, Chaojun
Huang, Guan
Zhao, Guosheng
Li, Haoyun
Zhu, Jiagang
Li, Kerui
Xu, Mengyuan
Deng, Qiuping
Wang, Siting
Qin, Wenkang
Chen, Xinze
Wang, Xiaofeng
Wang, Yankai
Cao, Yu
Chang, Yifan
Xu, Yuan
Ye, Yun
Wang, Yang
Zhou, Yukun
Zhang, Zhengyuan
Dong, Zhehao
Zhu, Zheng
contents World models are emerging as a foundational paradigm for scalable, data-efficient embodied AI. In this work, we present GigaWorld-0, a unified world model framework designed explicitly as a data engine for Vision-Language-Action (VLA) learning. GigaWorld-0 integrates two synergistic components: GigaWorld-0-Video, which leverages large-scale video generation to produce diverse, texture-rich, and temporally coherent embodied sequences under fine-grained control of appearance, camera viewpoint, and action semantics; and GigaWorld-0-3D, which combines 3D generative modeling, 3D Gaussian Splatting reconstruction, physically differentiable system identification, and executable motion planning to ensure geometric consistency and physical realism. Their joint optimization enables the scalable synthesis of embodied interaction data that is visually compelling, spatially coherent, physically plausible, and instruction-aligned. Training at scale is made feasible through our efficient GigaTrain framework, which exploits FP8-precision and sparse attention to drastically reduce memory and compute requirements. We conduct comprehensive evaluations showing that GigaWorld-0 generates high-quality, diverse, and controllable data across multiple dimensions. Critically, VLA model (e.g., GigaBrain-0) trained on GigaWorld-0-generated data achieve strong real-world performance, significantly improving generalization and task success on physical robots without any real-world interaction during training.
format Preprint
id arxiv_https___arxiv_org_abs_2511_19861
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle GigaWorld-0: World Models as Data Engine to Empower Embodied AI
GigaWorld Team
Ye, Angen
Wang, Boyuan
Ni, Chaojun
Huang, Guan
Zhao, Guosheng
Li, Haoyun
Zhu, Jiagang
Li, Kerui
Xu, Mengyuan
Deng, Qiuping
Wang, Siting
Qin, Wenkang
Chen, Xinze
Wang, Xiaofeng
Wang, Yankai
Cao, Yu
Chang, Yifan
Xu, Yuan
Ye, Yun
Wang, Yang
Zhou, Yukun
Zhang, Zhengyuan
Dong, Zhehao
Zhu, Zheng
Computer Vision and Pattern Recognition
Robotics
World models are emerging as a foundational paradigm for scalable, data-efficient embodied AI. In this work, we present GigaWorld-0, a unified world model framework designed explicitly as a data engine for Vision-Language-Action (VLA) learning. GigaWorld-0 integrates two synergistic components: GigaWorld-0-Video, which leverages large-scale video generation to produce diverse, texture-rich, and temporally coherent embodied sequences under fine-grained control of appearance, camera viewpoint, and action semantics; and GigaWorld-0-3D, which combines 3D generative modeling, 3D Gaussian Splatting reconstruction, physically differentiable system identification, and executable motion planning to ensure geometric consistency and physical realism. Their joint optimization enables the scalable synthesis of embodied interaction data that is visually compelling, spatially coherent, physically plausible, and instruction-aligned. Training at scale is made feasible through our efficient GigaTrain framework, which exploits FP8-precision and sparse attention to drastically reduce memory and compute requirements. We conduct comprehensive evaluations showing that GigaWorld-0 generates high-quality, diverse, and controllable data across multiple dimensions. Critically, VLA model (e.g., GigaBrain-0) trained on GigaWorld-0-generated data achieve strong real-world performance, significantly improving generalization and task success on physical robots without any real-world interaction during training.
title GigaWorld-0: World Models as Data Engine to Empower Embodied AI
topic Computer Vision and Pattern Recognition
Robotics
url https://arxiv.org/abs/2511.19861