DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , , , , , , , , , , , , , , , , , , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866914311380664320 |
|---|---|
| author | Gao, Shenyuan Liang, William Zheng, Kaiyuan Malik, Ayaan Ye, Seonghyeon Yu, Sihyun Tseng, Wei-Cheng Dong, Yuzhu Mo, Kaichun Lin, Chen-Hsuan Ma, Qianli Nah, Seungjun Magne, Loic Xiang, Jiannan Xie, Yuqi Zheng, Ruijie Niu, Dantong Tan, You Liang Zentner, K. R. Kurian, George Indupuru, Suneel Jannaty, Pooya Gu, Jinwei Zhang, Jun Malik, Jitendra Abbeel, Pieter Liu, Ming-Yu Zhu, Yuke Jang, Joel Fan, Linxi "Jim" |
| author_facet | Gao, Shenyuan Liang, William Zheng, Kaiyuan Malik, Ayaan Ye, Seonghyeon Yu, Sihyun Tseng, Wei-Cheng Dong, Yuzhu Mo, Kaichun Lin, Chen-Hsuan Ma, Qianli Nah, Seungjun Magne, Loic Xiang, Jiannan Xie, Yuqi Zheng, Ruijie Niu, Dantong Tan, You Liang Zentner, K. R. Kurian, George Indupuru, Suneel Jannaty, Pooya Gu, Jinwei Zhang, Jun Malik, Jitendra Abbeel, Pieter Liu, Ming-Yu Zhu, Yuke Jang, Joel Fan, Linxi "Jim" |
| contents | Being able to simulate the outcomes of actions in varied environments will revolutionize the development of generalist agents at scale. However, modeling these world dynamics, especially for dexterous robotics tasks, poses significant challenges due to limited data coverage and scarce action labels. As an endeavor towards this end, we introduce DreamDojo, a foundation world model that learns diverse interactions and dexterous controls from 44k hours of egocentric human videos. Our data mixture represents the largest video dataset to date for world model pretraining, spanning a wide range of daily scenarios with diverse objects and skills. To address the scarcity of action labels, we introduce continuous latent actions as unified proxy actions, enhancing interaction knowledge transfer from unlabeled videos. After post-training on small-scale target robot data, DreamDojo demonstrates a strong understanding of physics and precise action controllability. We also devise a distillation pipeline that accelerates DreamDojo to a real-time speed of 10.81 FPS and further improves context consistency. Our work enables several important applications based on generative world models, including live teleoperation, policy evaluation, and model-based planning. Systematic evaluation on multiple challenging out-of-distribution (OOD) benchmarks verifies the significance of our method for simulating open-world, contact-rich tasks, paving the way for general-purpose robot world models. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2602_06949 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos Gao, Shenyuan Liang, William Zheng, Kaiyuan Malik, Ayaan Ye, Seonghyeon Yu, Sihyun Tseng, Wei-Cheng Dong, Yuzhu Mo, Kaichun Lin, Chen-Hsuan Ma, Qianli Nah, Seungjun Magne, Loic Xiang, Jiannan Xie, Yuqi Zheng, Ruijie Niu, Dantong Tan, You Liang Zentner, K. R. Kurian, George Indupuru, Suneel Jannaty, Pooya Gu, Jinwei Zhang, Jun Malik, Jitendra Abbeel, Pieter Liu, Ming-Yu Zhu, Yuke Jang, Joel Fan, Linxi "Jim" Robotics Artificial Intelligence Computer Vision and Pattern Recognition Machine Learning Being able to simulate the outcomes of actions in varied environments will revolutionize the development of generalist agents at scale. However, modeling these world dynamics, especially for dexterous robotics tasks, poses significant challenges due to limited data coverage and scarce action labels. As an endeavor towards this end, we introduce DreamDojo, a foundation world model that learns diverse interactions and dexterous controls from 44k hours of egocentric human videos. Our data mixture represents the largest video dataset to date for world model pretraining, spanning a wide range of daily scenarios with diverse objects and skills. To address the scarcity of action labels, we introduce continuous latent actions as unified proxy actions, enhancing interaction knowledge transfer from unlabeled videos. After post-training on small-scale target robot data, DreamDojo demonstrates a strong understanding of physics and precise action controllability. We also devise a distillation pipeline that accelerates DreamDojo to a real-time speed of 10.81 FPS and further improves context consistency. Our work enables several important applications based on generative world models, including live teleoperation, policy evaluation, and model-based planning. Systematic evaluation on multiple challenging out-of-distribution (OOD) benchmarks verifies the significance of our method for simulating open-world, contact-rich tasks, paving the way for general-purpose robot world models. |
| title | DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos |
| topic | Robotics Artificial Intelligence Computer Vision and Pattern Recognition Machine Learning |
| url | https://arxiv.org/abs/2602.06949 |