Actions as Language: Fine-Tuning VLMs into VLAs Without Catastrophic Forgetting
Fuente:
arXiv
Saved in:
| Main Authors: | Hancock, Asher J., Wu, Xindi, Zha, Lihan, Russakovsky, Olga, Majumdar, Anirudha |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Run-time Observation Interventions Make Vision-Language-Action Models More Visually Robust
by: Hancock, Asher J., et al.
Published: (2024)
by: Hancock, Asher J., et al.
Published: (2024)
LAP: Language-Action Pre-Training Enables Zero-shot Cross-Embodiment Transfer
by: Zha, Lihan, et al.
Published: (2026)
by: Zha, Lihan, et al.
Published: (2026)
Video Generation Models in Robotics -- Applications, Research Challenges, Future Directions
by: Mei, Zhiting, et al.
Published: (2026)
by: Mei, Zhiting, et al.
Published: (2026)
How Do VLAs Effectively Inherit from VLMs?
by: Zhang, Chuheng, et al.
Published: (2025)
by: Zhang, Chuheng, et al.
Published: (2025)
Deceptive Risk Minimization: Out-of-Distribution Generalization by Deceiving Distribution Shift Detectors
by: Majumdar, Anirudha
Published: (2025)
by: Majumdar, Anirudha
Published: (2025)
Reliable and Scalable Robot Policy Evaluation with Imperfect Simulators
by: Badithela, Apurva, et al.
Published: (2025)
by: Badithela, Apurva, et al.
Published: (2025)
LEGS: Fine-Tuning Teleop-Free VLAs for Humanoid Loco-manipulation in an Embodied Gaussian Splatting World
by: Kim, Hojune, et al.
Published: (2026)
by: Kim, Hojune, et al.
Published: (2026)
Is Your Imitation Learning Policy Better than Mine? Policy Comparison with Near-Optimal Stopping
by: Snyder, David, et al.
Published: (2025)
by: Snyder, David, et al.
Published: (2025)
WoMAP: World Models For Embodied Open-Vocabulary Object Localization
by: Yin, Tenny, et al.
Published: (2025)
by: Yin, Tenny, et al.
Published: (2025)
Do World Action Models Generalize Better than VLAs? A Robustness Study
by: Zhang, Zhanguang, et al.
Published: (2026)
by: Zhang, Zhanguang, et al.
Published: (2026)
Guiding Data Collection via Factored Scaling Curves
by: Zha, Lihan, et al.
Published: (2025)
by: Zha, Lihan, et al.
Published: (2025)
PlayWorld: Learning Robot World Models from Autonomous Play
by: Yin, Tenny, et al.
Published: (2026)
by: Yin, Tenny, et al.
Published: (2026)
Predictive Red Teaming: Breaking Policies Without Breaking Robots
by: Majumdar, Anirudha, et al.
Published: (2025)
by: Majumdar, Anirudha, et al.
Published: (2025)
Privacy-Preserving Map-Free Exploration for Confirming the Absence of a Radioactive Source
by: Lepowsky, Eric, et al.
Published: (2024)
by: Lepowsky, Eric, et al.
Published: (2024)
Geometry Meets Vision: Revisiting Pretrained Semantics in Distilled Fields
by: Mei, Zhiting, et al.
Published: (2025)
by: Mei, Zhiting, et al.
Published: (2025)
Running VLAs at Real-time Speed
by: Ma, Yunchao, et al.
Published: (2025)
by: Ma, Yunchao, et al.
Published: (2025)
DoReMi: Grounding Language Model by Detecting and Recovering from Plan-Execution Misalignment
by: Guo, Yanjiang, et al.
Published: (2023)
by: Guo, Yanjiang, et al.
Published: (2023)
SITCOM: Scaling Inference-Time COMpute for VLAs
by: Saxena, Ayudh, et al.
Published: (2025)
by: Saxena, Ayudh, et al.
Published: (2025)
How VLAs (Really) Work In Open-World Environments
by: Rasouli, Amir, et al.
Published: (2026)
by: Rasouli, Amir, et al.
Published: (2026)
VLAs are Confined yet Capable of Generalizing to Novel Instructions
by: Li, Quanyi
Published: (2025)
by: Li, Quanyi
Published: (2025)
Shallow-π: Knowledge Distillation for Flow-based VLAs
by: Jeon, Boseong, et al.
Published: (2026)
by: Jeon, Boseong, et al.
Published: (2026)
Primitive Subspaces Mediate Few-Shot Transfer in VLAs
by: Singh, Anya, et al.
Published: (2026)
by: Singh, Anya, et al.
Published: (2026)
Vision-Language Dataset Distillation
by: Wu, Xindi, et al.
Published: (2023)
by: Wu, Xindi, et al.
Published: (2023)
SciFi-Benchmark: Leveraging Science Fiction To Improve Robot Behavior
by: Sermanet, Pierre, et al.
Published: (2025)
by: Sermanet, Pierre, et al.
Published: (2025)
SIREN: Semantic, Initialization-Free Registration of Multi-Robot Gaussian Splatting Maps
by: Shorinwa, Ola, et al.
Published: (2025)
by: Shorinwa, Ola, et al.
Published: (2025)
STARE-VLA: Progressive Stage-Aware Reinforcement for Fine-Tuning Vision-Language-Action Models
by: Xu, Feng, et al.
Published: (2025)
by: Xu, Feng, et al.
Published: (2025)
Reinforcement Fine-Tuning of Flow-Matching Policies for Vision-Language-Action Models
by: Lyu, Mingyang, et al.
Published: (2025)
by: Lyu, Mingyang, et al.
Published: (2025)
When Vision Overrides Language: Evaluating and Mitigating Counterfactual Failures in VLAs
by: Fang, Yu, et al.
Published: (2026)
by: Fang, Yu, et al.
Published: (2026)
Visual Compositional Tuning
by: Wu, Xindi, et al.
Published: (2025)
by: Wu, Xindi, et al.
Published: (2025)
mimic-video: Video-Action Models for Generalizable Robot Control Beyond VLAs
by: Pai, Jonas, et al.
Published: (2025)
by: Pai, Jonas, et al.
Published: (2025)
VTAM: Video-Tactile-Action Models for Complex Physical Interaction Beyond VLAs
by: Yuan, Haoran, et al.
Published: (2026)
by: Yuan, Haoran, et al.
Published: (2026)
Notes-to-Self: Scratchpad Augmented VLAs for Memory Dependent Manipulation Tasks
by: Haresh, Sanjay, et al.
Published: (2026)
by: Haresh, Sanjay, et al.
Published: (2026)
cVLA: Towards Efficient Camera-Space VLAs
by: Argus, Max, et al.
Published: (2025)
by: Argus, Max, et al.
Published: (2025)
Beyond Objects: Contextual Synthetic Data Generation for Fine-Grained Classification
by: Yang, William, et al.
Published: (2025)
by: Yang, William, et al.
Published: (2025)
FASTER: Rethinking Real-Time Flow VLAs
by: Lu, Yuxiang, et al.
Published: (2026)
by: Lu, Yuxiang, et al.
Published: (2026)
Realtime-VLA V2: Learning to Run VLAs Fast, Smooth, and Accurate
by: Yang, Chen, et al.
Published: (2026)
by: Yang, Chen, et al.
Published: (2026)
VLA-0: Building State-of-the-Art VLAs with Zero Modification
by: Goyal, Ankit, et al.
Published: (2025)
by: Goyal, Ankit, et al.
Published: (2025)
Risk-Calibrated Human-Robot Interaction via Set-Valued Intent Prediction
by: Lidard, Justin, et al.
Published: (2024)
by: Lidard, Justin, et al.
Published: (2024)
Differentiate-and-Inject: Enhancing VLAs via Functional Differentiation Induced by In-Parameter Structural Reasoning
by: Hou, Jingyi, et al.
Published: (2026)
by: Hou, Jingyi, et al.
Published: (2026)
Physically Grounded Vision-Language Models for Robotic Manipulation
by: Gao, Jensen, et al.
Published: (2023)
by: Gao, Jensen, et al.
Published: (2023)
Similar Items
-
Run-time Observation Interventions Make Vision-Language-Action Models More Visually Robust
by: Hancock, Asher J., et al.
Published: (2024) -
LAP: Language-Action Pre-Training Enables Zero-shot Cross-Embodiment Transfer
by: Zha, Lihan, et al.
Published: (2026) -
Video Generation Models in Robotics -- Applications, Research Challenges, Future Directions
by: Mei, Zhiting, et al.
Published: (2026) -
How Do VLAs Effectively Inherit from VLMs?
by: Zhang, Chuheng, et al.
Published: (2025) -
Deceptive Risk Minimization: Out-of-Distribution Generalization by Deceiving Distribution Shift Detectors
by: Majumdar, Anirudha
Published: (2025)