Turning Video Models into Generalist Robot Policies
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Sizhe Lester, Kim, Evan, Bai, Xingjian, Zhao, Tong, Pang, Tao, Simchowitz, Max, Sitzmann, Vincent |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Diffusion Forcing: Next-token Prediction Meets Full-Sequence Diffusion
by: Chen, Boyuan, et al.
Published: (2024)
by: Chen, Boyuan, et al.
Published: (2024)
Shortcut Learning in Generalist Robot Policies: The Role of Dataset Diversity and Fragmentation
by: Xing, Youguang, et al.
Published: (2025)
by: Xing, Youguang, et al.
Published: (2025)
Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations
by: Hu, Yucheng, et al.
Published: (2024)
by: Hu, Yucheng, et al.
Published: (2024)
Controlling diverse robots by inferring Jacobian fields with deep networks
by: Li, Sizhe Lester, et al.
Published: (2024)
by: Li, Sizhe Lester, et al.
Published: (2024)
InternVLA-M1: A Spatially Guided Vision-Language-Action Framework for Generalist Robot Policy
by: Chen, Xinyi, et al.
Published: (2025)
by: Chen, Xinyi, et al.
Published: (2025)
DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos
by: Gao, Shenyuan, et al.
Published: (2026)
by: Gao, Shenyuan, et al.
Published: (2026)
History-Guided Video Diffusion
by: Song, Kiwhan, et al.
Published: (2025)
by: Song, Kiwhan, et al.
Published: (2025)
Vidar: Embodied Video Diffusion Model for Generalist Manipulation
by: Feng, Yao, et al.
Published: (2025)
by: Feng, Yao, et al.
Published: (2025)
On-Device Diffusion Transformer Policy for Efficient Robot Manipulation
by: Wu, Yiming, et al.
Published: (2025)
by: Wu, Yiming, et al.
Published: (2025)
Large Video Planner Enables Generalizable Robot Control
by: Chen, Boyuan, et al.
Published: (2025)
by: Chen, Boyuan, et al.
Published: (2025)
DRAW2ACT: Turning Depth-Encoded Trajectories into Robotic Demonstration Videos
by: Bai, Yang, et al.
Published: (2025)
by: Bai, Yang, et al.
Published: (2025)
Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy
by: Hou, Zhi, et al.
Published: (2025)
by: Hou, Zhi, et al.
Published: (2025)
What Matters in Building Vision-Language-Action Models for Generalist Robots
by: Li, Xinghang, et al.
Published: (2024)
by: Li, Xinghang, et al.
Published: (2024)
Scaling View Synthesis Transformers
by: Kim, Evan, et al.
Published: (2026)
by: Kim, Evan, et al.
Published: (2026)
OctoNav: Towards Generalist Embodied Navigation
by: Gao, Chen, et al.
Published: (2025)
by: Gao, Chen, et al.
Published: (2025)
GeRM: A Generalist Robotic Model with Mixture-of-experts for Quadruped Robot
by: Song, Wenxuan, et al.
Published: (2024)
by: Song, Wenxuan, et al.
Published: (2024)
ReFineVLA: Multimodal Reasoning-Aware Generalist Robotic Policies via Teacher-Guided Fine-Tuning
by: Van Vo, Tuan, et al.
Published: (2026)
by: Van Vo, Tuan, et al.
Published: (2026)
GLEAM: Learning Generalizable Exploration Policy for Active Mapping in Complex 3D Indoor Scenes
by: Chen, Xiao, et al.
Published: (2025)
by: Chen, Xiao, et al.
Published: (2025)
The One RING: a Robotic Indoor Navigation Generalist
by: Eftekhar, Ainaz, et al.
Published: (2024)
by: Eftekhar, Ainaz, et al.
Published: (2024)
SurgSora: Object-Aware Diffusion Model for Controllable Surgical Video Generation
by: Chen, Tong, et al.
Published: (2024)
by: Chen, Tong, et al.
Published: (2024)
NORA: A Small Open-Sourced Generalist Vision Language Action Model for Embodied Tasks
by: Hung, Chia-Yu, et al.
Published: (2025)
by: Hung, Chia-Yu, et al.
Published: (2025)
VG4D: Vision-Language Model Goes 4D Video Recognition
by: Deng, Zhichao, et al.
Published: (2024)
by: Deng, Zhichao, et al.
Published: (2024)
RoboVIP: Multi-View Video Generation with Visual Identity Prompting Augments Robot Manipulation
by: Wang, Boyang, et al.
Published: (2026)
by: Wang, Boyang, et al.
Published: (2026)
Spatial Policy: Guiding Visuomotor Robotic Manipulation with Spatial-Aware Modeling and Reasoning
by: Liu, Yijun, et al.
Published: (2025)
by: Liu, Yijun, et al.
Published: (2025)
VideoVLA: Video Generators Can Be Generalizable Robot Manipulators
by: Shen, Yichao, et al.
Published: (2025)
by: Shen, Yichao, et al.
Published: (2025)
Explainable Adversarial-Robust Vision-Language-Action Model for Robotic Manipulation
by: Kim, Ju-Young, et al.
Published: (2025)
by: Kim, Ju-Young, et al.
Published: (2025)
Goal Force: Teaching Video Models To Accomplish Physics-Conditioned Goals
by: Gillman, Nate, et al.
Published: (2026)
by: Gillman, Nate, et al.
Published: (2026)
TapSampling: Inference-Time Sampling with a Task-Progress-Understanding Verifier for Robotic Manipulation
by: Zhao, Sizhe, et al.
Published: (2026)
by: Zhao, Sizhe, et al.
Published: (2026)
Robotic Programmer: Video Instructed Policy Code Generation for Robotic Manipulation
by: Xie, Senwei, et al.
Published: (2025)
by: Xie, Senwei, et al.
Published: (2025)
PianoMime: Learning a Generalist, Dexterous Piano Player from Internet Demonstrations
by: Qian, Cheng, et al.
Published: (2024)
by: Qian, Cheng, et al.
Published: (2024)
Cosmos-H-Surgical: Learning Surgical Robot Policies from Videos via World Modeling
by: He, Yufan, et al.
Published: (2025)
by: He, Yufan, et al.
Published: (2025)
Novel View Synthesis with Neural Radiance Fields for Industrial Robot Applications
by: Hillemann, Markus, et al.
Published: (2024)
by: Hillemann, Markus, et al.
Published: (2024)
Robots Pre-train Robots: Manipulation-Centric Robotic Representation from Large-Scale Robot Datasets
by: Jiang, Guangqi, et al.
Published: (2024)
by: Jiang, Guangqi, et al.
Published: (2024)
This&That: Language-Gesture Controlled Video Generation for Robot Planning
by: Wang, Boyang, et al.
Published: (2024)
by: Wang, Boyang, et al.
Published: (2024)
Fast Visuomotor Policy for Robotic Manipulation
by: Jia, Jingkai, et al.
Published: (2025)
by: Jia, Jingkai, et al.
Published: (2025)
IRASim: A Fine-Grained World Model for Robot Manipulation
by: Zhu, Fangqi, et al.
Published: (2024)
by: Zhu, Fangqi, et al.
Published: (2024)
Efficient Robotic Policy Learning via Latent Space Backward Planning
by: Liu, Dongxiu, et al.
Published: (2025)
by: Liu, Dongxiu, et al.
Published: (2025)
GenNBV: Generalizable Next-Best-View Policy for Active 3D Reconstruction
by: Chen, Xiao, et al.
Published: (2024)
by: Chen, Xiao, et al.
Published: (2024)
Robotic Manipulation by Imitating Generated Videos Without Physical Demonstrations
by: Patel, Shivansh, et al.
Published: (2025)
by: Patel, Shivansh, et al.
Published: (2025)
Robotic Visual Instruction
by: Li, Yanbang, et al.
Published: (2025)
by: Li, Yanbang, et al.
Published: (2025)
Similar Items
-
Diffusion Forcing: Next-token Prediction Meets Full-Sequence Diffusion
by: Chen, Boyuan, et al.
Published: (2024) -
Shortcut Learning in Generalist Robot Policies: The Role of Dataset Diversity and Fragmentation
by: Xing, Youguang, et al.
Published: (2025) -
Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations
by: Hu, Yucheng, et al.
Published: (2024) -
Controlling diverse robots by inferring Jacobian fields with deep networks
by: Li, Sizhe Lester, et al.
Published: (2024) -
InternVLA-M1: A Spatially Guided Vision-Language-Action Framework for Generalist Robot Policy
by: Chen, Xinyi, et al.
Published: (2025)