Turning Video Models into Generalist Robot Policies

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Sizhe Lester, Kim, Evan, Bai, Xingjian, Zhao, Tong, Pang, Tao, Simchowitz, Max, Sitzmann, Vincent
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918526403477504
author Li, Sizhe Lester
Kim, Evan
Bai, Xingjian
Zhao, Tong
Pang, Tao
Simchowitz, Max
Sitzmann, Vincent
author_facet Li, Sizhe Lester
Kim, Evan
Bai, Xingjian
Zhao, Tong
Pang, Tao
Simchowitz, Max
Sitzmann, Vincent
contents Video generative models have emerged as a promising robotics backbone, capable of generating videos that depict the completion of complex tasks across embodiments and environments. Recent work proposes robot foundation models that jointly predict future observations and actions by finetuning video models with action-labeled data. In this paper, we test the limits of an alternative approach: leave the video planner as-is while training an embodiment-specific inverse dynamics model (IDM). This decoupling offers several natural benefits: the video planner remains embodiment-agnostic, different video models can be interchanged easily without re-training the IDM, and the IDM can be independently trained with readily available self-play data. We present a closed-loop, video-to-action policy that combines an action-free video world model with a carefully-designed IDM based on the robot embodiment Jacobian. We demonstrate that our IDM design is both data-efficient and scalable to high-dimensional action spaces. Our policy, which we coin the Video-to-Embodied Robot Action Model (VERA), achieves strong performance across simulated and real-world benchmarks, including zero-shot Panda arm manipulation and 16-DoF Allegro-hand dexterous cube re-orientation. The same video planner can be used across multiple embodiments by pairing it with different embodiment-specific IDMs. Our results show that decoupled video planning plus faithful video-to-action translation is a viable alternative route towards zero-shot, cross-embodiment, and generalizable robot control. More results are available on our project website: https://vera.csail.mit.edu.
format Preprint
id arxiv_https___arxiv_org_abs_2605_27817
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Turning Video Models into Generalist Robot Policies
Li, Sizhe Lester
Kim, Evan
Bai, Xingjian
Zhao, Tong
Pang, Tao
Simchowitz, Max
Sitzmann, Vincent
Robotics
Artificial Intelligence
Computer Vision and Pattern Recognition
Machine Learning
Video generative models have emerged as a promising robotics backbone, capable of generating videos that depict the completion of complex tasks across embodiments and environments. Recent work proposes robot foundation models that jointly predict future observations and actions by finetuning video models with action-labeled data. In this paper, we test the limits of an alternative approach: leave the video planner as-is while training an embodiment-specific inverse dynamics model (IDM). This decoupling offers several natural benefits: the video planner remains embodiment-agnostic, different video models can be interchanged easily without re-training the IDM, and the IDM can be independently trained with readily available self-play data. We present a closed-loop, video-to-action policy that combines an action-free video world model with a carefully-designed IDM based on the robot embodiment Jacobian. We demonstrate that our IDM design is both data-efficient and scalable to high-dimensional action spaces. Our policy, which we coin the Video-to-Embodied Robot Action Model (VERA), achieves strong performance across simulated and real-world benchmarks, including zero-shot Panda arm manipulation and 16-DoF Allegro-hand dexterous cube re-orientation. The same video planner can be used across multiple embodiments by pairing it with different embodiment-specific IDMs. Our results show that decoupled video planning plus faithful video-to-action translation is a viable alternative route towards zero-shot, cross-embodiment, and generalizable robot control. More results are available on our project website: https://vera.csail.mit.edu.
title Turning Video Models into Generalist Robot Policies
topic Robotics
Artificial Intelligence
Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2605.27817