DRAW2ACT: Turning Depth-Encoded Trajectories into Robotic Demonstration Videos

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bai, Yang, Yang, Liudi, Eskandar, George, Shen, Fengyi, Altillawi, Mohammad, Liu, Ziyuan, Kutyniok, Gitta
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917148462415872
author Bai, Yang
Yang, Liudi
Eskandar, George
Shen, Fengyi
Altillawi, Mohammad
Liu, Ziyuan
Kutyniok, Gitta
author_facet Bai, Yang
Yang, Liudi
Eskandar, George
Shen, Fengyi
Altillawi, Mohammad
Liu, Ziyuan
Kutyniok, Gitta
contents Video diffusion models provide powerful real-world simulators for embodied AI but remain limited in controllability for robotic manipulation. Recent works on trajectory-conditioned video generation address this gap but often rely on 2D trajectories or single modality conditioning, which restricts their ability to produce controllable and consistent robotic demonstrations. We present DRAW2ACT, a depth-aware trajectory-conditioned video generation framework that extracts multiple orthogonal representations from the input trajectory, capturing depth, semantics, shape and motion, and injects them into the diffusion model. Moreover, we propose to jointly generate spatially aligned RGB and depth videos, leveraging cross-modality attention mechanisms and depth supervision to enhance the spatio-temporal consistency. Finally, we introduce a multimodal policy model conditioned on the generated RGB and depth sequences to regress the robot's joint angles. Experiments on Bridge V2, Berkeley Autolab, and simulation benchmarks show that DRAW2ACT achieves superior visual fidelity and consistency while yielding higher manipulation success rates compared to existing baselines.
format Preprint
id arxiv_https___arxiv_org_abs_2512_14217
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DRAW2ACT: Turning Depth-Encoded Trajectories into Robotic Demonstration Videos
Bai, Yang
Yang, Liudi
Eskandar, George
Shen, Fengyi
Altillawi, Mohammad
Liu, Ziyuan
Kutyniok, Gitta
Computer Vision and Pattern Recognition
Robotics
Video diffusion models provide powerful real-world simulators for embodied AI but remain limited in controllability for robotic manipulation. Recent works on trajectory-conditioned video generation address this gap but often rely on 2D trajectories or single modality conditioning, which restricts their ability to produce controllable and consistent robotic demonstrations. We present DRAW2ACT, a depth-aware trajectory-conditioned video generation framework that extracts multiple orthogonal representations from the input trajectory, capturing depth, semantics, shape and motion, and injects them into the diffusion model. Moreover, we propose to jointly generate spatially aligned RGB and depth videos, leveraging cross-modality attention mechanisms and depth supervision to enhance the spatio-temporal consistency. Finally, we introduce a multimodal policy model conditioned on the generated RGB and depth sequences to regress the robot's joint angles. Experiments on Bridge V2, Berkeley Autolab, and simulation benchmarks show that DRAW2ACT achieves superior visual fidelity and consistency while yielding higher manipulation success rates compared to existing baselines.
title DRAW2ACT: Turning Depth-Encoded Trajectories into Robotic Demonstration Videos
topic Computer Vision and Pattern Recognition
Robotics
url https://arxiv.org/abs/2512.14217