World Simulation with Video Foundation Models for Physical AI
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866915815123582976 |
|---|---|
| author | NVIDIA : Ali, Arslan Bai, Junjie Bala, Maciej Balaji, Yogesh Blakeman, Aaron Cai, Tiffany Cao, Jiaxin Cao, Tianshi Cha, Elizabeth Chao, Yu-Wei Chattopadhyay, Prithvijit Chen, Mike Chen, Yongxin Chen, Yu Cheng, Shuai Cui, Yin Diamond, Jenna Ding, Yifan Fan, Jiaojiao Fan, Linxi Feng, Liang Ferroni, Francesco Fidler, Sanja Fu, Xiao Gao, Ruiyuan Ge, Yunhao Gu, Jinwei Gupta, Aryaman Gururani, Siddharth Hanafi, Imad El Hassani, Ali Hao, Zekun Huffman, Jacob Jang, Joel Jannaty, Pooya Kautz, Jan Lam, Grace Li, Xuan Li, Zhaoshuo Liao, Maosheng Lin, Chen-Hsuan Lin, Tsung-Yi Lin, Yen-Chen Ling, Huan Liu, Ming-Yu Liu, Xian Lu, Yifan Luo, Alice Ma, Qianli Mao, Hanzi Mo, Kaichun Nah, Seungjun Narang, Yashraj Panaskar, Abhijeet Pavao, Lindsey Pham, Trung Ramezanali, Morteza Reda, Fitsum Reed, Scott Ren, Xuanchi Shao, Haonan Shen, Yue Shi, Stella Song, Shuran Stefaniak, Bartosz Sun, Shangkun Tang, Shitao Tasmeen, Sameena Tchapmi, Lyne Tseng, Wei-Cheng Varghese, Jibin Wang, Andrew Z. Wang, Hao Wang, Haoxiang Wang, Heng Wang, Ting-Chun Wei, Fangyin Xu, Jiashu Yang, Dinghao Yang, Xiaodong Ye, Haotian Ye, Seonghyeon Zeng, Xiaohui Zhang, Jing Zhang, Qinsheng Zheng, Kaiwen Zhu, Andrew Zhu, Yuke |
| author_facet | NVIDIA : Ali, Arslan Bai, Junjie Bala, Maciej Balaji, Yogesh Blakeman, Aaron Cai, Tiffany Cao, Jiaxin Cao, Tianshi Cha, Elizabeth Chao, Yu-Wei Chattopadhyay, Prithvijit Chen, Mike Chen, Yongxin Chen, Yu Cheng, Shuai Cui, Yin Diamond, Jenna Ding, Yifan Fan, Jiaojiao Fan, Linxi Feng, Liang Ferroni, Francesco Fidler, Sanja Fu, Xiao Gao, Ruiyuan Ge, Yunhao Gu, Jinwei Gupta, Aryaman Gururani, Siddharth Hanafi, Imad El Hassani, Ali Hao, Zekun Huffman, Jacob Jang, Joel Jannaty, Pooya Kautz, Jan Lam, Grace Li, Xuan Li, Zhaoshuo Liao, Maosheng Lin, Chen-Hsuan Lin, Tsung-Yi Lin, Yen-Chen Ling, Huan Liu, Ming-Yu Liu, Xian Lu, Yifan Luo, Alice Ma, Qianli Mao, Hanzi Mo, Kaichun Nah, Seungjun Narang, Yashraj Panaskar, Abhijeet Pavao, Lindsey Pham, Trung Ramezanali, Morteza Reda, Fitsum Reed, Scott Ren, Xuanchi Shao, Haonan Shen, Yue Shi, Stella Song, Shuran Stefaniak, Bartosz Sun, Shangkun Tang, Shitao Tasmeen, Sameena Tchapmi, Lyne Tseng, Wei-Cheng Varghese, Jibin Wang, Andrew Z. Wang, Hao Wang, Haoxiang Wang, Heng Wang, Ting-Chun Wei, Fangyin Xu, Jiashu Yang, Dinghao Yang, Xiaodong Ye, Haotian Ye, Seonghyeon Zeng, Xiaohui Zhang, Jing Zhang, Qinsheng Zheng, Kaiwen Zhu, Andrew Zhu, Yuke |
| contents | We introduce [Cosmos-Predict2.5], the latest generation of the Cosmos World Foundation Models for Physical AI. Built on a flow-based architecture, [Cosmos-Predict2.5] unifies Text2World, Image2World, and Video2World generation in a single model and leverages [Cosmos-Reason1], a Physical AI vision-language model, to provide richer text grounding and finer control of world simulation. Trained on 200M curated video clips and refined with reinforcement learning-based post-training, [Cosmos-Predict2.5] achieves substantial improvements over [Cosmos-Predict1] in video quality and instruction alignment, with models released at 2B and 14B scales. These capabilities enable more reliable synthetic data generation, policy evaluation, and closed-loop simulation for robotics and autonomous systems. We further extend the family with [Cosmos-Transfer2.5], a control-net style framework for Sim2Real and Real2Real world translation. Despite being 3.5$\times$ smaller than [Cosmos-Transfer1], it delivers higher fidelity and robust long-horizon video generation. Together, these advances establish [Cosmos-Predict2.5] and [Cosmos-Transfer2.5] as versatile tools for scaling embodied intelligence. To accelerate research and deployment in Physical AI, we release source code, pretrained checkpoints, and curated benchmarks under the NVIDIA Open Model License at https://github.com/nvidia-cosmos/cosmos-predict2.5 and https://github.com/nvidia-cosmos/cosmos-transfer2.5. We hope these open resources lower the barrier to adoption and foster innovation in building the next generation of embodied intelligence. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2511_00062 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | World Simulation with Video Foundation Models for Physical AI NVIDIA : Ali, Arslan Bai, Junjie Bala, Maciej Balaji, Yogesh Blakeman, Aaron Cai, Tiffany Cao, Jiaxin Cao, Tianshi Cha, Elizabeth Chao, Yu-Wei Chattopadhyay, Prithvijit Chen, Mike Chen, Yongxin Chen, Yu Cheng, Shuai Cui, Yin Diamond, Jenna Ding, Yifan Fan, Jiaojiao Fan, Linxi Feng, Liang Ferroni, Francesco Fidler, Sanja Fu, Xiao Gao, Ruiyuan Ge, Yunhao Gu, Jinwei Gupta, Aryaman Gururani, Siddharth Hanafi, Imad El Hassani, Ali Hao, Zekun Huffman, Jacob Jang, Joel Jannaty, Pooya Kautz, Jan Lam, Grace Li, Xuan Li, Zhaoshuo Liao, Maosheng Lin, Chen-Hsuan Lin, Tsung-Yi Lin, Yen-Chen Ling, Huan Liu, Ming-Yu Liu, Xian Lu, Yifan Luo, Alice Ma, Qianli Mao, Hanzi Mo, Kaichun Nah, Seungjun Narang, Yashraj Panaskar, Abhijeet Pavao, Lindsey Pham, Trung Ramezanali, Morteza Reda, Fitsum Reed, Scott Ren, Xuanchi Shao, Haonan Shen, Yue Shi, Stella Song, Shuran Stefaniak, Bartosz Sun, Shangkun Tang, Shitao Tasmeen, Sameena Tchapmi, Lyne Tseng, Wei-Cheng Varghese, Jibin Wang, Andrew Z. Wang, Hao Wang, Haoxiang Wang, Heng Wang, Ting-Chun Wei, Fangyin Xu, Jiashu Yang, Dinghao Yang, Xiaodong Ye, Haotian Ye, Seonghyeon Zeng, Xiaohui Zhang, Jing Zhang, Qinsheng Zheng, Kaiwen Zhu, Andrew Zhu, Yuke Computer Vision and Pattern Recognition Artificial Intelligence Machine Learning Robotics We introduce [Cosmos-Predict2.5], the latest generation of the Cosmos World Foundation Models for Physical AI. Built on a flow-based architecture, [Cosmos-Predict2.5] unifies Text2World, Image2World, and Video2World generation in a single model and leverages [Cosmos-Reason1], a Physical AI vision-language model, to provide richer text grounding and finer control of world simulation. Trained on 200M curated video clips and refined with reinforcement learning-based post-training, [Cosmos-Predict2.5] achieves substantial improvements over [Cosmos-Predict1] in video quality and instruction alignment, with models released at 2B and 14B scales. These capabilities enable more reliable synthetic data generation, policy evaluation, and closed-loop simulation for robotics and autonomous systems. We further extend the family with [Cosmos-Transfer2.5], a control-net style framework for Sim2Real and Real2Real world translation. Despite being 3.5$\times$ smaller than [Cosmos-Transfer1], it delivers higher fidelity and robust long-horizon video generation. Together, these advances establish [Cosmos-Predict2.5] and [Cosmos-Transfer2.5] as versatile tools for scaling embodied intelligence. To accelerate research and deployment in Physical AI, we release source code, pretrained checkpoints, and curated benchmarks under the NVIDIA Open Model License at https://github.com/nvidia-cosmos/cosmos-predict2.5 and https://github.com/nvidia-cosmos/cosmos-transfer2.5. We hope these open resources lower the barrier to adoption and foster innovation in building the next generation of embodied intelligence. |
| title | World Simulation with Video Foundation Models for Physical AI |
| topic | Computer Vision and Pattern Recognition Artificial Intelligence Machine Learning Robotics |
| url | https://arxiv.org/abs/2511.00062 |