Knowledge Insulating Vision-Language-Action Models: Train Fast, Run Fast, Generalize Better
Fuente:
arXiv
Saved in:
| Main Authors: | Driess, Danny, Springenberg, Jost Tobias, Ichter, Brian, Yu, Lili, Li-Bell, Adrian, Pertsch, Karl, Ren, Allen Z., Walke, Homer, Vuong, Quan, Shi, Lucy Xiaoyang, Levine, Sergey |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MEM: Multi-Scale Embodied Memory for Vision Language Action Models
by: Torne, Marcel, et al.
Published: (2026)
by: Torne, Marcel, et al.
Published: (2026)
FAST: Efficient Action Tokenization for Vision-Language-Action Models
by: Pertsch, Karl, et al.
Published: (2025)
by: Pertsch, Karl, et al.
Published: (2025)
Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models
by: Shi, Lucy Xiaoyang, et al.
Published: (2025)
by: Shi, Lucy Xiaoyang, et al.
Published: (2025)
$π_{0.5}$: a Vision-Language-Action Model with Open-World Generalization
by: Intelligence, Physical, et al.
Published: (2025)
by: Intelligence, Physical, et al.
Published: (2025)
Training Strategies for Efficient Embodied Reasoning
by: Chen, William, et al.
Published: (2025)
by: Chen, William, et al.
Published: (2025)
$π_0$: A Vision-Language-Action Flow Model for General Robot Control
by: Black, Kevin, et al.
Published: (2024)
by: Black, Kevin, et al.
Published: (2024)
Steerable Vision-Language-Action Policies for Embodied Reasoning and Hierarchical Control
by: Chen, William, et al.
Published: (2026)
by: Chen, William, et al.
Published: (2026)
RL Token: Bootstrapping Online RL with Vision-Language-Action Models
by: Xu, Charles, et al.
Published: (2026)
by: Xu, Charles, et al.
Published: (2026)
Scaling Cross-Embodied Learning: One Policy for Manipulation, Navigation, Locomotion and Aviation
by: Doshi, Ria, et al.
Published: (2024)
by: Doshi, Ria, et al.
Published: (2024)
Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved)
by: Qin, Chongli, et al.
Published: (2025)
by: Qin, Chongli, et al.
Published: (2025)
Robust Finetuning of Vision-Language-Action Robot Policies via Parameter Merging
by: Yadav, Yajat, et al.
Published: (2025)
by: Yadav, Yajat, et al.
Published: (2025)
Autonomous Improvement of Instruction Following Skills via Foundation Models
by: Zhou, Zhiyuan, et al.
Published: (2024)
by: Zhou, Zhiyuan, et al.
Published: (2024)
KALIE: Fine-Tuning Vision-Language Models for Open-World Manipulation without Robot Data
by: Tang, Grace, et al.
Published: (2024)
by: Tang, Grace, et al.
Published: (2024)
Evaluating Real-World Robot Manipulation Policies in Simulation
by: Li, Xuanlin, et al.
Published: (2024)
by: Li, Xuanlin, et al.
Published: (2024)
Stabilizing Contrastive RL: Techniques for Robotic Goal Reaching from Offline Data
by: Zheng, Chongyi, et al.
Published: (2023)
by: Zheng, Chongyi, et al.
Published: (2023)
Yell At Your Robot: Improving On-the-Fly from Language Corrections
by: Shi, Lucy Xiaoyang, et al.
Published: (2024)
by: Shi, Lucy Xiaoyang, et al.
Published: (2024)
Octo: An Open-Source Generalist Robot Policy
by: Octo Model Team, et al.
Published: (2024)
by: Octo Model Team, et al.
Published: (2024)
Emergence of Human to Robot Transfer in Vision-Language-Action Models
by: Kareer, Simar, et al.
Published: (2025)
by: Kareer, Simar, et al.
Published: (2025)
RACER: Epistemic Risk-Sensitive RL Enables Fast Driving with Fewer Crashes
by: Stachowicz, Kyle, et al.
Published: (2024)
by: Stachowicz, Kyle, et al.
Published: (2024)
AutoEval: Autonomous Evaluation of Generalist Robot Manipulation Policies in the Real World
by: Zhou, Zhiyuan, et al.
Published: (2025)
by: Zhou, Zhiyuan, et al.
Published: (2025)
Running Fast in Ethiopia
by: Jordan, Robert
Published: (1972)
by: Jordan, Robert
Published: (1972)
RoboReward: General-Purpose Vision-Language Reward Models for Robotics
by: Lee, Tony, et al.
Published: (2026)
by: Lee, Tony, et al.
Published: (2026)
Robotic Control via Embodied Chain-of-Thought Reasoning
by: Zawalski, Michał, et al.
Published: (2024)
by: Zawalski, Michał, et al.
Published: (2024)
Game On: Towards Language Models as RL Experimenters
by: Zhang, Jingwei, et al.
Published: (2024)
by: Zhang, Jingwei, et al.
Published: (2024)
BridgeData V2: A Dataset for Robot Learning at Scale
by: Walke, Homer, et al.
Published: (2023)
by: Walke, Homer, et al.
Published: (2023)
SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities
by: Chen, Boyuan, et al.
Published: (2024)
by: Chen, Boyuan, et al.
Published: (2024)
AsyncVLA: An Asynchronous VLA for Fast and Robust Navigation on the Edge
by: Hirose, Noriaki, et al.
Published: (2026)
by: Hirose, Noriaki, et al.
Published: (2026)
OpenVLA: An Open-Source Vision-Language-Action Model
by: Kim, Moo Jin, et al.
Published: (2024)
by: Kim, Moo Jin, et al.
Published: (2024)
Goal Representations for Instruction Following: A Semi-Supervised Language Interface to Control
by: Myers, Vivek, et al.
Published: (2023)
by: Myers, Vivek, et al.
Published: (2023)
The Special Librarian as a Front Runner: Running Fast, Running Hard, Running Ahead.
by: Regan, Muriel
Published: (1990)
by: Regan, Muriel
Published: (1990)
Training-Time Action Conditioning for Efficient Real-Time Chunking
by: Black, Kevin, et al.
Published: (2025)
by: Black, Kevin, et al.
Published: (2025)
$π^{*}_{0.6}$: a VLA That Learns From Experience
by: Intelligence, Physical, et al.
Published: (2025)
by: Intelligence, Physical, et al.
Published: (2025)
Fast parallel sampling under isoperimetry
by: Anari, Nima, et al.
Published: (2024)
by: Anari, Nima, et al.
Published: (2024)
Thinking Forward and Backward: Effective Backward Planning with Large Language Models
by: Ren, Allen Z., et al.
Published: (2024)
by: Ren, Allen Z., et al.
Published: (2024)
Reinforcement Learning with Action Chunking
by: Li, Qiyang, et al.
Published: (2025)
by: Li, Qiyang, et al.
Published: (2025)
Fast Makespan Minimization via Short ILPs
by: Hermelin, Danny, et al.
Published: (2026)
by: Hermelin, Danny, et al.
Published: (2026)
CHEMICAL CONTROL OVER THE SYNTHESIS OF DEFECT-FREE AND HETEROMETAL-DOPED METAL OXIDE NANOPARTICLES BY THE SINGLE-SOURCE PRECURSOR APPROACH
by: M. Driess
Published: (2005)
by: M. Driess
Published: (2005)
PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs
by: Nasiriany, Soroush, et al.
Published: (2024)
by: Nasiriany, Soroush, et al.
Published: (2024)
FastStair: Learning to Run Up Stairs with Humanoid Robots
by: Liu, Yan, et al.
Published: (2026)
by: Liu, Yan, et al.
Published: (2026)
Fast and Certifiable Trajectory Optimization
by: Kang, Shucheng, et al.
Published: (2024)
by: Kang, Shucheng, et al.
Published: (2024)
Similar Items
-
MEM: Multi-Scale Embodied Memory for Vision Language Action Models
by: Torne, Marcel, et al.
Published: (2026) -
FAST: Efficient Action Tokenization for Vision-Language-Action Models
by: Pertsch, Karl, et al.
Published: (2025) -
Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models
by: Shi, Lucy Xiaoyang, et al.
Published: (2025) -
$π_{0.5}$: a Vision-Language-Action Model with Open-World Generalization
by: Intelligence, Physical, et al.
Published: (2025) -
Training Strategies for Efficient Embodied Reasoning
by: Chen, William, et al.
Published: (2025)