VideoVLA: Video Generators Can Be Generalizable Robot Manipulators

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shen, Yichao, Wei, Fangyun, Du, Zhiying, Liang, Yaobo, Lu, Yan, Yang, Jiaolong, Zheng, Nanning, Guo, Baining
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909947984347136
author Shen, Yichao
Wei, Fangyun
Du, Zhiying
Liang, Yaobo
Lu, Yan
Yang, Jiaolong
Zheng, Nanning
Guo, Baining
author_facet Shen, Yichao
Wei, Fangyun
Du, Zhiying
Liang, Yaobo
Lu, Yan
Yang, Jiaolong
Zheng, Nanning
Guo, Baining
contents Generalization in robot manipulation is essential for deploying robots in open-world environments and advancing toward artificial general intelligence. While recent Vision-Language-Action (VLA) models leverage large pre-trained understanding models for perception and instruction following, their ability to generalize to novel tasks, objects, and settings remains limited. In this work, we present VideoVLA, a simple approach that explores the potential of transforming large video generation models into robotic VLA manipulators. Given a language instruction and an image, VideoVLA predicts an action sequence as well as the future visual outcomes. Built on a multi-modal Diffusion Transformer, VideoVLA jointly models video, language, and action modalities, using pre-trained video generative models for joint visual and action forecasting. Our experiments show that high-quality imagined futures correlate with reliable action predictions and task success, highlighting the importance of visual imagination in manipulation. VideoVLA demonstrates strong generalization, including imitating other embodiments' skills and handling novel objects. This dual-prediction strategy - forecasting both actions and their visual consequences - explores a paradigm shift in robot learning and unlocks generalization capabilities in manipulation systems.
format Preprint
id arxiv_https___arxiv_org_abs_2512_06963
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle VideoVLA: Video Generators Can Be Generalizable Robot Manipulators
Shen, Yichao
Wei, Fangyun
Du, Zhiying
Liang, Yaobo
Lu, Yan
Yang, Jiaolong
Zheng, Nanning
Guo, Baining
Robotics
Artificial Intelligence
Computer Vision and Pattern Recognition
Generalization in robot manipulation is essential for deploying robots in open-world environments and advancing toward artificial general intelligence. While recent Vision-Language-Action (VLA) models leverage large pre-trained understanding models for perception and instruction following, their ability to generalize to novel tasks, objects, and settings remains limited. In this work, we present VideoVLA, a simple approach that explores the potential of transforming large video generation models into robotic VLA manipulators. Given a language instruction and an image, VideoVLA predicts an action sequence as well as the future visual outcomes. Built on a multi-modal Diffusion Transformer, VideoVLA jointly models video, language, and action modalities, using pre-trained video generative models for joint visual and action forecasting. Our experiments show that high-quality imagined futures correlate with reliable action predictions and task success, highlighting the importance of visual imagination in manipulation. VideoVLA demonstrates strong generalization, including imitating other embodiments' skills and handling novel objects. This dual-prediction strategy - forecasting both actions and their visual consequences - explores a paradigm shift in robot learning and unlocks generalization capabilities in manipulation systems.
title VideoVLA: Video Generators Can Be Generalizable Robot Manipulators
topic Robotics
Artificial Intelligence
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.06963