D2E: Scaling Vision-Action Pretraining on Desktop Data for Transfer to Embodied AI

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Choi, Suhwan, Jung, Jaeyoon, Seong, Haebin, Kim, Minchan, Kim, Minyeong, Cho, Yongjun, Kim, Yoonshik, Park, Yubeen, Yu, Youngjae, Lee, Yunsung
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866912937943236608
author Choi, Suhwan
Jung, Jaeyoon
Seong, Haebin
Kim, Minchan
Kim, Minyeong
Cho, Yongjun
Kim, Yoonshik
Park, Yubeen
Yu, Youngjae
Lee, Yunsung
author_facet Choi, Suhwan
Jung, Jaeyoon
Seong, Haebin
Kim, Minchan
Kim, Minyeong
Cho, Yongjun
Kim, Yoonshik
Park, Yubeen
Yu, Youngjae
Lee, Yunsung
contents Large language models leverage internet-scale text data, yet embodied AI remains constrained by the prohibitive costs of physical trajectory collection. Desktop environments -- particularly gaming -- offer a compelling alternative: they provide rich sensorimotor interactions at scale while maintaining the structured observation-action coupling essential for embodied learning. We present D2E (Desktop to Embodied AI), a framework that demonstrates desktop interactions can serve as an effective pretraining substrate for robotics embodied AI tasks. Unlike prior work that remained domain-specific (e.g., VPT for Minecraft) or kept data proprietary (e.g., SIMA), D2E establishes a complete pipeline from scalable desktop data collection to verified transfer in embodied domains. Our framework comprises three components: (1) the OWA Toolkit that unifies diverse desktop interactions into a standardized format with 152x compression, (2) the Generalist-IDM that achieves strong zero-shot generalization across unseen games through timestamp-based event prediction, enabling internet-scale pseudo-labeling, and (3) VAPT that transfers desktop-pretrained representations to physical manipulation and navigation. Using 1.3K+ hours of data (259 hours of human demonstrations and 1K+ hours of pseudo-labeled gameplay), our 1B-parameter model achieves 96.6% success on LIBERO manipulation and 83.3% on CANVAS navigation, matching or surpassing models up to 7x larger, such as π_{0} (3.3B) and OpenVLA (7B). These results demonstrate that sensorimotor primitives learned from digital interactions transfer effectively to real-world physical tasks, establishing desktop pretraining as a practical paradigm for embodied AI. All resources are publicly available at https://worv-ai.github.io/d2e.
format Preprint
id arxiv_https___arxiv_org_abs_2510_05684
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle D2E: Scaling Vision-Action Pretraining on Desktop Data for Transfer to Embodied AI
Choi, Suhwan
Jung, Jaeyoon
Seong, Haebin
Kim, Minchan
Kim, Minyeong
Cho, Yongjun
Kim, Yoonshik
Park, Yubeen
Yu, Youngjae
Lee, Yunsung
Artificial Intelligence
Computer Vision and Pattern Recognition
Robotics
Large language models leverage internet-scale text data, yet embodied AI remains constrained by the prohibitive costs of physical trajectory collection. Desktop environments -- particularly gaming -- offer a compelling alternative: they provide rich sensorimotor interactions at scale while maintaining the structured observation-action coupling essential for embodied learning. We present D2E (Desktop to Embodied AI), a framework that demonstrates desktop interactions can serve as an effective pretraining substrate for robotics embodied AI tasks. Unlike prior work that remained domain-specific (e.g., VPT for Minecraft) or kept data proprietary (e.g., SIMA), D2E establishes a complete pipeline from scalable desktop data collection to verified transfer in embodied domains. Our framework comprises three components: (1) the OWA Toolkit that unifies diverse desktop interactions into a standardized format with 152x compression, (2) the Generalist-IDM that achieves strong zero-shot generalization across unseen games through timestamp-based event prediction, enabling internet-scale pseudo-labeling, and (3) VAPT that transfers desktop-pretrained representations to physical manipulation and navigation. Using 1.3K+ hours of data (259 hours of human demonstrations and 1K+ hours of pseudo-labeled gameplay), our 1B-parameter model achieves 96.6% success on LIBERO manipulation and 83.3% on CANVAS navigation, matching or surpassing models up to 7x larger, such as π_{0} (3.3B) and OpenVLA (7B). These results demonstrate that sensorimotor primitives learned from digital interactions transfer effectively to real-world physical tasks, establishing desktop pretraining as a practical paradigm for embodied AI. All resources are publicly available at https://worv-ai.github.io/d2e.
title D2E: Scaling Vision-Action Pretraining on Desktop Data for Transfer to Embodied AI
topic Artificial Intelligence
Computer Vision and Pattern Recognition
Robotics
url https://arxiv.org/abs/2510.05684