RoboEnvision: A Long-Horizon Video Generation Model for Multi-Task Robot Manipulation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Liudi, Bai, Yang, Eskandar, George, Shen, Fengyi, Altillawi, Mohammad, Chen, Dong, Majumder, Soumajit, Liu, Ziyuan, Kutyniok, Gitta, Valada, Abhinav
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916812807995392
author Yang, Liudi
Bai, Yang
Eskandar, George
Shen, Fengyi
Altillawi, Mohammad
Chen, Dong
Majumder, Soumajit
Liu, Ziyuan
Kutyniok, Gitta
Valada, Abhinav
author_facet Yang, Liudi
Bai, Yang
Eskandar, George
Shen, Fengyi
Altillawi, Mohammad
Chen, Dong
Majumder, Soumajit
Liu, Ziyuan
Kutyniok, Gitta
Valada, Abhinav
contents We address the problem of generating long-horizon videos for robotic manipulation tasks. Text-to-video diffusion models have made significant progress in photorealism, language understanding, and motion generation but struggle with long-horizon robotic tasks. Recent works use video diffusion models for high-quality simulation data and predictive rollouts in robot planning. However, these works predict short sequences of the robot achieving one task and employ an autoregressive paradigm to extend to the long horizon, leading to error accumulations in the generated video and in the execution. To overcome these limitations, we propose a novel pipeline that bypasses the need for autoregressive generation. We achieve this through a threefold contribution: 1) we first decompose the high-level goals into smaller atomic tasks and generate keyframes aligned with these instructions. A second diffusion model then interpolates between each of the two generated frames, achieving the long-horizon video. 2) We propose a semantics preserving attention module to maintain consistency between the keyframes. 3) We design a lightweight policy model to regress the robot joint states from generated videos. Our approach achieves state-of-the-art results on two benchmarks in video quality and consistency while outperforming previous policy models on long-horizon tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2506_22007
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle RoboEnvision: A Long-Horizon Video Generation Model for Multi-Task Robot Manipulation
Yang, Liudi
Bai, Yang
Eskandar, George
Shen, Fengyi
Altillawi, Mohammad
Chen, Dong
Majumder, Soumajit
Liu, Ziyuan
Kutyniok, Gitta
Valada, Abhinav
Computer Vision and Pattern Recognition
We address the problem of generating long-horizon videos for robotic manipulation tasks. Text-to-video diffusion models have made significant progress in photorealism, language understanding, and motion generation but struggle with long-horizon robotic tasks. Recent works use video diffusion models for high-quality simulation data and predictive rollouts in robot planning. However, these works predict short sequences of the robot achieving one task and employ an autoregressive paradigm to extend to the long horizon, leading to error accumulations in the generated video and in the execution. To overcome these limitations, we propose a novel pipeline that bypasses the need for autoregressive generation. We achieve this through a threefold contribution: 1) we first decompose the high-level goals into smaller atomic tasks and generate keyframes aligned with these instructions. A second diffusion model then interpolates between each of the two generated frames, achieving the long-horizon video. 2) We propose a semantics preserving attention module to maintain consistency between the keyframes. 3) We design a lightweight policy model to regress the robot joint states from generated videos. Our approach achieves state-of-the-art results on two benchmarks in video quality and consistency while outperforming previous policy models on long-horizon tasks.
title RoboEnvision: A Long-Horizon Video Generation Model for Multi-Task Robot Manipulation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.22007