V-Dreamer: Automating Robotic Simulation and Trajectory Synthesis via Video Generation Priors

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: He, Songjia, Chen, Zixuan, Ding, Hongyu, Shao, Dian, Shi, Jieqi, Li, Chenxu, Huo, Jing, Gao, Yang
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866918397957111808
author He, Songjia
Chen, Zixuan
Ding, Hongyu
Shao, Dian
Shi, Jieqi
Li, Chenxu
Huo, Jing
Gao, Yang
author_facet He, Songjia
Chen, Zixuan
Ding, Hongyu
Shao, Dian
Shi, Jieqi
Li, Chenxu
Huo, Jing
Gao, Yang
contents Training generalist robots demands large-scale, diverse manipulation data, yet real-world collection is prohibitively expensive, and existing simulators are often constrained by fixed asset libraries and manual heuristics. To bridge this gap, we present V-Dreamer, a fully automated framework that generates open-vocabulary, simulation-ready manipulation environments and executable expert trajectories directly from natural language instructions. V-Dreamer employs a novel generative pipeline that constructs physically grounded 3D scenes using large language models and 3D generative models, validated by geometric constraints to ensure stable, collision-free layouts. Crucially, for behavior synthesis, we leverage video generation models as rich motion priors. These visual predictions are then mapped into executable robot trajectories via a robust Sim-to-Gen visual-kinematic alignment module utilizing CoTracker3 and VGGT. This pipeline supports high visual diversity and physical fidelity without manual intervention. To evaluate the generated data, we train imitation learning policies on synthesized trajectories encompassing diverse object and environment variations. Extensive evaluations on tabletop manipulation tasks using the Piper robotic arm demonstrate that our policies robustly generalize to unseen objects in simulation and achieve effective sim-to-real transfer, successfully manipulating novel real-world objects.
format Preprint
id arxiv_https___arxiv_org_abs_2603_18811
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle V-Dreamer: Automating Robotic Simulation and Trajectory Synthesis via Video Generation Priors
He, Songjia
Chen, Zixuan
Ding, Hongyu
Shao, Dian
Shi, Jieqi
Li, Chenxu
Huo, Jing
Gao, Yang
Robotics
Training generalist robots demands large-scale, diverse manipulation data, yet real-world collection is prohibitively expensive, and existing simulators are often constrained by fixed asset libraries and manual heuristics. To bridge this gap, we present V-Dreamer, a fully automated framework that generates open-vocabulary, simulation-ready manipulation environments and executable expert trajectories directly from natural language instructions. V-Dreamer employs a novel generative pipeline that constructs physically grounded 3D scenes using large language models and 3D generative models, validated by geometric constraints to ensure stable, collision-free layouts. Crucially, for behavior synthesis, we leverage video generation models as rich motion priors. These visual predictions are then mapped into executable robot trajectories via a robust Sim-to-Gen visual-kinematic alignment module utilizing CoTracker3 and VGGT. This pipeline supports high visual diversity and physical fidelity without manual intervention. To evaluate the generated data, we train imitation learning policies on synthesized trajectories encompassing diverse object and environment variations. Extensive evaluations on tabletop manipulation tasks using the Piper robotic arm demonstrate that our policies robustly generalize to unseen objects in simulation and achieve effective sim-to-real transfer, successfully manipulating novel real-world objects.
title V-Dreamer: Automating Robotic Simulation and Trajectory Synthesis via Video Generation Priors
topic Robotics
url https://arxiv.org/abs/2603.18811