Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations
Fuente:
arXiv
Saved in:
| Main Authors: | Hu, Yucheng, Guo, Yanjiang, Wang, Pengchao, Chen, Xiaoyu, Wang, Yen-Jen, Zhang, Jianke, Sreenath, Koushil, Lu, Chaochao, Chen, Jianyu |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Prediction with Action: Visual Policy Learning via Joint Denoising Process
by: Guo, Yanjiang, et al.
Published: (2024)
by: Guo, Yanjiang, et al.
Published: (2024)
UniJEPA: Enhancing Robot Policy via Unified Continuous and Discrete Representation Learning
by: Zhang, Jianke, et al.
Published: (2025)
by: Zhang, Jianke, et al.
Published: (2025)
HiRT: Enhancing Robotic Control with Hierarchical Robot Transformers
by: Zhang, Jianke, et al.
Published: (2024)
by: Zhang, Jianke, et al.
Published: (2024)
Improving Vision-Language-Action Model with Online Reinforcement Learning
by: Guo, Yanjiang, et al.
Published: (2025)
by: Guo, Yanjiang, et al.
Published: (2025)
Prompt a Robot to Walk with Large Language Models
by: Wang, Yen-Jen, et al.
Published: (2023)
by: Wang, Yen-Jen, et al.
Published: (2023)
Veo-Act: How Far Can Frontier Video Models Advance Generalizable Robot Manipulation?
by: Zhang, Zhongru, et al.
Published: (2026)
by: Zhang, Zhongru, et al.
Published: (2026)
UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent
by: Zhang, Jianke, et al.
Published: (2025)
by: Zhang, Jianke, et al.
Published: (2025)
Robust Adversarial Policy Optimization Under Dynamics Uncertainty
by: Kim, Mintae, et al.
Published: (2026)
by: Kim, Mintae, et al.
Published: (2026)
DoReMi: Grounding Language Model by Detecting and Recovering from Plan-Execution Misalignment
by: Guo, Yanjiang, et al.
Published: (2023)
by: Guo, Yanjiang, et al.
Published: (2023)
Coordinated Humanoid Manipulation with Choice Policies
by: Qi, Haozhi, et al.
Published: (2025)
by: Qi, Haozhi, et al.
Published: (2025)
Turning Video Models into Generalist Robot Policies
by: Li, Sizhe Lester, et al.
Published: (2026)
by: Li, Sizhe Lester, et al.
Published: (2026)
Robotic Programmer: Video Instructed Policy Code Generation for Robotic Manipulation
by: Xie, Senwei, et al.
Published: (2025)
by: Xie, Senwei, et al.
Published: (2025)
DDAT: Diffusion Policies Enforcing Dynamically Admissible Robot Trajectories
by: Bouvier, Jean-Baptiste, et al.
Published: (2025)
by: Bouvier, Jean-Baptiste, et al.
Published: (2025)
BagelVLA: Enhancing Long-Horizon Manipulation via Interleaved Vision-Language-Action Generation
by: Hu, Yucheng, et al.
Published: (2026)
by: Hu, Yucheng, et al.
Published: (2026)
Learning Smooth Humanoid Locomotion through Lipschitz-Constrained Policies
by: Chen, Zixuan, et al.
Published: (2024)
by: Chen, Zixuan, et al.
Published: (2024)
Humanoid Locomotion as Next Token Prediction
by: Radosavovic, Ilija, et al.
Published: (2024)
by: Radosavovic, Ilija, et al.
Published: (2024)
MomaGraph: State-Aware Unified Scene Graphs with Vision-Language Model for Embodied Task Planning
by: Ju, Yuanchen, et al.
Published: (2025)
by: Ju, Yuanchen, et al.
Published: (2025)
NanoVLA: Routing Decoupled Vision-Language Understanding for Nano-sized Generalist Robotic Policies
by: Chen, Jiahong, et al.
Published: (2025)
by: Chen, Jiahong, et al.
Published: (2025)
InternVLA-M1: A Spatially Guided Vision-Language-Action Framework for Generalist Robot Policy
by: Chen, Xinyi, et al.
Published: (2025)
by: Chen, Xinyi, et al.
Published: (2025)
DINOv3-Diffusion Policy: Self-Supervised Large Visual Model for Visuomotor Diffusion Policy Learning
by: Egbe, ThankGod, et al.
Published: (2025)
by: Egbe, ThankGod, et al.
Published: (2025)
Shortcut Learning in Generalist Robot Policies: The Role of Dataset Diversity and Fragmentation
by: Xing, Youguang, et al.
Published: (2025)
by: Xing, Youguang, et al.
Published: (2025)
VLAW: Iterative Co-Improvement of Vision-Language-Action Policy and World Model
by: Guo, Yanjiang, et al.
Published: (2026)
by: Guo, Yanjiang, et al.
Published: (2026)
Towards a Generalizable Bimanual Foundation Policy via Flow-based Video Prediction
by: Fan, Chenyou, et al.
Published: (2025)
by: Fan, Chenyou, et al.
Published: (2025)
Decentralized Navigation of a Cable-Towed Load using Quadrupedal Robot Team via MARL
by: Chen, Wen-Tse, et al.
Published: (2025)
by: Chen, Wen-Tse, et al.
Published: (2025)
Learning Generalizable Robot Policy with Human Demonstration Video as a Prompt
by: Zhu, Xiang, et al.
Published: (2025)
by: Zhu, Xiang, et al.
Published: (2025)
Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy
by: Hou, Zhi, et al.
Published: (2025)
by: Hou, Zhi, et al.
Published: (2025)
WOMBET: World Model-based Experience Transfer for Robust and Sample-efficient Reinforcement Learning
by: Kim, Mintae, et al.
Published: (2026)
by: Kim, Mintae, et al.
Published: (2026)
Humanoid-Gym: Reinforcement Learning for Humanoid Robot with Zero-Shot Sim2Real Transfer
by: Gu, Xinyang, et al.
Published: (2024)
by: Gu, Xinyang, et al.
Published: (2024)
Learning Dexterous Manipulation Skills from Imperfect Simulations
by: Hsieh, Elvis, et al.
Published: (2025)
by: Hsieh, Elvis, et al.
Published: (2025)
FRAPPE: Infusing World Modeling into Generalist Policies via Multiple Future Representation Alignment
by: Zhao, Han, et al.
Published: (2026)
by: Zhao, Han, et al.
Published: (2026)
ReliOcc: Towards Reliable Semantic Occupancy Prediction via Uncertainty Learning
by: Wang, Song, et al.
Published: (2024)
by: Wang, Song, et al.
Published: (2024)
Ctrl-World: A Controllable Generative World Model for Robot Manipulation
by: Guo, Yanjiang, et al.
Published: (2025)
by: Guo, Yanjiang, et al.
Published: (2025)
Learning-based Trajectory Tracking for Bird-inspired Flapping-Wing Robots
by: Cai, Jiaze, et al.
Published: (2024)
by: Cai, Jiaze, et al.
Published: (2024)
Advancing Humanoid Locomotion: Mastering Challenging Terrains with Denoising World Model Learning
by: Gu, Xinyang, et al.
Published: (2024)
by: Gu, Xinyang, et al.
Published: (2024)
Grounding Bodily Awareness in Visual Representations for Efficient Policy Learning
by: Wang, Junlin, et al.
Published: (2025)
by: Wang, Junlin, et al.
Published: (2025)
ReFineVLA: Multimodal Reasoning-Aware Generalist Robotic Policies via Teacher-Guided Fine-Tuning
by: Van Vo, Tuan, et al.
Published: (2026)
by: Van Vo, Tuan, et al.
Published: (2026)
Octo: An Open-Source Generalist Robot Policy
by: Octo Model Team, et al.
Published: (2024)
by: Octo Model Team, et al.
Published: (2024)
SPARK: Skeleton-Parameter Aligned Retargeting on Humanoid Robots with Kinodynamic Trajectory Optimization
by: Wang, Hanwen, et al.
Published: (2026)
by: Wang, Hanwen, et al.
Published: (2026)
UAM: A Dual-Stream Perspective on Forgetting in VLA Training
by: Zhang, Jianke, et al.
Published: (2026)
by: Zhang, Jianke, et al.
Published: (2026)
Fast Visuomotor Policy for Robotic Manipulation
by: Jia, Jingkai, et al.
Published: (2025)
by: Jia, Jingkai, et al.
Published: (2025)
Similar Items
-
Prediction with Action: Visual Policy Learning via Joint Denoising Process
by: Guo, Yanjiang, et al.
Published: (2024) -
UniJEPA: Enhancing Robot Policy via Unified Continuous and Discrete Representation Learning
by: Zhang, Jianke, et al.
Published: (2025) -
HiRT: Enhancing Robotic Control with Hierarchical Robot Transformers
by: Zhang, Jianke, et al.
Published: (2024) -
Improving Vision-Language-Action Model with Online Reinforcement Learning
by: Guo, Yanjiang, et al.
Published: (2025) -
Prompt a Robot to Walk with Large Language Models
by: Wang, Yen-Jen, et al.
Published: (2023)