AnchorDream: Repurposing Video Diffusion for Embodiment-Aware Robot Data Synthesis

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Ye, Junjie, Xue, Rong, Van Hoorick, Basile, Tokmakov, Pavel, Irshad, Muhammad Zubair, Wang, Yue, Guizilini, Vitor
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912760375279616
author Ye, Junjie
Xue, Rong
Van Hoorick, Basile
Tokmakov, Pavel
Irshad, Muhammad Zubair
Wang, Yue
Guizilini, Vitor
author_facet Ye, Junjie
Xue, Rong
Van Hoorick, Basile
Tokmakov, Pavel
Irshad, Muhammad Zubair
Wang, Yue
Guizilini, Vitor
contents The collection of large-scale and diverse robot demonstrations remains a major bottleneck for imitation learning, as real-world data acquisition is costly and simulators offer limited diversity and fidelity with pronounced sim-to-real gaps. While generative models present an attractive solution, existing methods often alter only visual appearances without creating new behaviors, or suffer from embodiment inconsistencies that yield implausible motions. To address these limitations, we introduce AnchorDream, an embodiment-aware world model that repurposes pretrained video diffusion models for robot data synthesis. AnchorDream conditions the diffusion process on robot motion renderings, anchoring the embodiment to prevent hallucination while synthesizing objects and environments consistent with the robot's kinematics. Starting from only a handful of human teleoperation demonstrations, our method scales them into large, diverse, high-quality datasets without requiring explicit environment modeling. Experiments show that the generated data leads to consistent improvements in downstream policy learning, with relative gains of 36.4% in simulator benchmarks and nearly double performance in real-world studies. These results suggest that grounding generative world models in robot motion provides a practical path toward scaling imitation learning.
format Preprint
id arxiv_https___arxiv_org_abs_2512_11797
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle AnchorDream: Repurposing Video Diffusion for Embodiment-Aware Robot Data Synthesis
Ye, Junjie
Xue, Rong
Van Hoorick, Basile
Tokmakov, Pavel
Irshad, Muhammad Zubair
Wang, Yue
Guizilini, Vitor
Robotics
Computer Vision and Pattern Recognition
The collection of large-scale and diverse robot demonstrations remains a major bottleneck for imitation learning, as real-world data acquisition is costly and simulators offer limited diversity and fidelity with pronounced sim-to-real gaps. While generative models present an attractive solution, existing methods often alter only visual appearances without creating new behaviors, or suffer from embodiment inconsistencies that yield implausible motions. To address these limitations, we introduce AnchorDream, an embodiment-aware world model that repurposes pretrained video diffusion models for robot data synthesis. AnchorDream conditions the diffusion process on robot motion renderings, anchoring the embodiment to prevent hallucination while synthesizing objects and environments consistent with the robot's kinematics. Starting from only a handful of human teleoperation demonstrations, our method scales them into large, diverse, high-quality datasets without requiring explicit environment modeling. Experiments show that the generated data leads to consistent improvements in downstream policy learning, with relative gains of 36.4% in simulator benchmarks and nearly double performance in real-world studies. These results suggest that grounding generative world models in robot motion provides a practical path toward scaling imitation learning.
title AnchorDream: Repurposing Video Diffusion for Embodiment-Aware Robot Data Synthesis
topic Robotics
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.11797