DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Gao, Shenyuan, Liang, William, Zheng, Kaiyuan, Malik, Ayaan, Ye, Seonghyeon, Yu, Sihyun, Tseng, Wei-Cheng, Dong, Yuzhu, Mo, Kaichun, Lin, Chen-Hsuan, Ma, Qianli, Nah, Seungjun, Magne, Loic, Xiang, Jiannan, Xie, Yuqi, Zheng, Ruijie, Niu, Dantong, Tan, You Liang, Zentner, K. R., Kurian, George, Indupuru, Suneel, Jannaty, Pooya, Gu, Jinwei, Zhang, Jun, Malik, Jitendra, Abbeel, Pieter, Liu, Ming-Yu, Zhu, Yuke, Jang, Joel, Fan, Linxi "Jim"
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914311380664320
author Gao, Shenyuan
Liang, William
Zheng, Kaiyuan
Malik, Ayaan
Ye, Seonghyeon
Yu, Sihyun
Tseng, Wei-Cheng
Dong, Yuzhu
Mo, Kaichun
Lin, Chen-Hsuan
Ma, Qianli
Nah, Seungjun
Magne, Loic
Xiang, Jiannan
Xie, Yuqi
Zheng, Ruijie
Niu, Dantong
Tan, You Liang
Zentner, K. R.
Kurian, George
Indupuru, Suneel
Jannaty, Pooya
Gu, Jinwei
Zhang, Jun
Malik, Jitendra
Abbeel, Pieter
Liu, Ming-Yu
Zhu, Yuke
Jang, Joel
Fan, Linxi "Jim"
author_facet Gao, Shenyuan
Liang, William
Zheng, Kaiyuan
Malik, Ayaan
Ye, Seonghyeon
Yu, Sihyun
Tseng, Wei-Cheng
Dong, Yuzhu
Mo, Kaichun
Lin, Chen-Hsuan
Ma, Qianli
Nah, Seungjun
Magne, Loic
Xiang, Jiannan
Xie, Yuqi
Zheng, Ruijie
Niu, Dantong
Tan, You Liang
Zentner, K. R.
Kurian, George
Indupuru, Suneel
Jannaty, Pooya
Gu, Jinwei
Zhang, Jun
Malik, Jitendra
Abbeel, Pieter
Liu, Ming-Yu
Zhu, Yuke
Jang, Joel
Fan, Linxi "Jim"
contents Being able to simulate the outcomes of actions in varied environments will revolutionize the development of generalist agents at scale. However, modeling these world dynamics, especially for dexterous robotics tasks, poses significant challenges due to limited data coverage and scarce action labels. As an endeavor towards this end, we introduce DreamDojo, a foundation world model that learns diverse interactions and dexterous controls from 44k hours of egocentric human videos. Our data mixture represents the largest video dataset to date for world model pretraining, spanning a wide range of daily scenarios with diverse objects and skills. To address the scarcity of action labels, we introduce continuous latent actions as unified proxy actions, enhancing interaction knowledge transfer from unlabeled videos. After post-training on small-scale target robot data, DreamDojo demonstrates a strong understanding of physics and precise action controllability. We also devise a distillation pipeline that accelerates DreamDojo to a real-time speed of 10.81 FPS and further improves context consistency. Our work enables several important applications based on generative world models, including live teleoperation, policy evaluation, and model-based planning. Systematic evaluation on multiple challenging out-of-distribution (OOD) benchmarks verifies the significance of our method for simulating open-world, contact-rich tasks, paving the way for general-purpose robot world models.
format Preprint
id arxiv_https___arxiv_org_abs_2602_06949
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos
Gao, Shenyuan
Liang, William
Zheng, Kaiyuan
Malik, Ayaan
Ye, Seonghyeon
Yu, Sihyun
Tseng, Wei-Cheng
Dong, Yuzhu
Mo, Kaichun
Lin, Chen-Hsuan
Ma, Qianli
Nah, Seungjun
Magne, Loic
Xiang, Jiannan
Xie, Yuqi
Zheng, Ruijie
Niu, Dantong
Tan, You Liang
Zentner, K. R.
Kurian, George
Indupuru, Suneel
Jannaty, Pooya
Gu, Jinwei
Zhang, Jun
Malik, Jitendra
Abbeel, Pieter
Liu, Ming-Yu
Zhu, Yuke
Jang, Joel
Fan, Linxi "Jim"
Robotics
Artificial Intelligence
Computer Vision and Pattern Recognition
Machine Learning
Being able to simulate the outcomes of actions in varied environments will revolutionize the development of generalist agents at scale. However, modeling these world dynamics, especially for dexterous robotics tasks, poses significant challenges due to limited data coverage and scarce action labels. As an endeavor towards this end, we introduce DreamDojo, a foundation world model that learns diverse interactions and dexterous controls from 44k hours of egocentric human videos. Our data mixture represents the largest video dataset to date for world model pretraining, spanning a wide range of daily scenarios with diverse objects and skills. To address the scarcity of action labels, we introduce continuous latent actions as unified proxy actions, enhancing interaction knowledge transfer from unlabeled videos. After post-training on small-scale target robot data, DreamDojo demonstrates a strong understanding of physics and precise action controllability. We also devise a distillation pipeline that accelerates DreamDojo to a real-time speed of 10.81 FPS and further improves context consistency. Our work enables several important applications based on generative world models, including live teleoperation, policy evaluation, and model-based planning. Systematic evaluation on multiple challenging out-of-distribution (OOD) benchmarks verifies the significance of our method for simulating open-world, contact-rich tasks, paving the way for general-purpose robot world models.
title DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos
topic Robotics
Artificial Intelligence
Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2602.06949