Steerable Vision-Language-Action Policies for Embodied Reasoning and Hierarchical Control
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, William, Bhatia, Jagdeep Singh, Glossop, Catherine, Mathihalli, Nikhil, Doshi, Ria, Tang, Andy, Driess, Danny, Pertsch, Karl, Levine, Sergey |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Training Strategies for Efficient Embodied Reasoning
by: Chen, William, et al.
Published: (2025)
by: Chen, William, et al.
Published: (2025)
FAST: Efficient Action Tokenization for Vision-Language-Action Models
by: Pertsch, Karl, et al.
Published: (2025)
by: Pertsch, Karl, et al.
Published: (2025)
OmniVLA: An Omni-Modal Vision-Language-Action Model for Robot Navigation
by: Hirose, Noriaki, et al.
Published: (2025)
by: Hirose, Noriaki, et al.
Published: (2025)
Scaling Cross-Embodied Learning: One Policy for Manipulation, Navigation, Locomotion and Aviation
by: Doshi, Ria, et al.
Published: (2024)
by: Doshi, Ria, et al.
Published: (2024)
CAST: Counterfactual Labels Improve Instruction Following in Vision-Language-Action Models
by: Glossop, Catherine, et al.
Published: (2025)
by: Glossop, Catherine, et al.
Published: (2025)
Robust Finetuning of Vision-Language-Action Robot Policies via Parameter Merging
by: Yadav, Yajat, et al.
Published: (2025)
by: Yadav, Yajat, et al.
Published: (2025)
Robotic Control via Embodied Chain-of-Thought Reasoning
by: Zawalski, Michał, et al.
Published: (2024)
by: Zawalski, Michał, et al.
Published: (2024)
MEM: Multi-Scale Embodied Memory for Vision Language Action Models
by: Torne, Marcel, et al.
Published: (2026)
by: Torne, Marcel, et al.
Published: (2026)
Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models
by: Shi, Lucy Xiaoyang, et al.
Published: (2025)
by: Shi, Lucy Xiaoyang, et al.
Published: (2025)
Knowledge Insulating Vision-Language-Action Models: Train Fast, Run Fast, Generalize Better
by: Driess, Danny, et al.
Published: (2025)
by: Driess, Danny, et al.
Published: (2025)
AsyncVLA: An Asynchronous VLA for Fast and Robust Navigation on the Edge
by: Hirose, Noriaki, et al.
Published: (2026)
by: Hirose, Noriaki, et al.
Published: (2026)
LeLaN: Learning A Language-Conditioned Navigation Policy from In-the-Wild Videos
by: Hirose, Noriaki, et al.
Published: (2024)
by: Hirose, Noriaki, et al.
Published: (2024)
Emergence of Human to Robot Transfer in Vision-Language-Action Models
by: Kareer, Simar, et al.
Published: (2025)
by: Kareer, Simar, et al.
Published: (2025)
AutoEval: Autonomous Evaluation of Generalist Robot Manipulation Policies in the Real World
by: Zhou, Zhiyuan, et al.
Published: (2025)
by: Zhou, Zhiyuan, et al.
Published: (2025)
RoboReward: General-Purpose Vision-Language Reward Models for Robotics
by: Lee, Tony, et al.
Published: (2026)
by: Lee, Tony, et al.
Published: (2026)
Shadow: Leveraging Segmentation Masks for Cross-Embodiment Policy Transfer
by: Lepert, Marion, et al.
Published: (2025)
by: Lepert, Marion, et al.
Published: (2025)
SteerVLA: Steering Vision-Language-Action Models in Long-Tail Driving Scenarios
by: Gao, Tian, et al.
Published: (2026)
by: Gao, Tian, et al.
Published: (2026)
Learning to Drive Anywhere with Model-Based Reannotation
by: Hirose, Noriaki, et al.
Published: (2025)
by: Hirose, Noriaki, et al.
Published: (2025)
$k$-loose elements and $k$-paving matroids
by: Singh, Jagdeep
Published: (2024)
by: Singh, Jagdeep
Published: (2024)
VAMOS: A Hierarchical Vision-Language-Action Model for Capability-Modulated and Steerable Navigation
by: Castro, Mateo Guaman, et al.
Published: (2025)
by: Castro, Mateo Guaman, et al.
Published: (2025)
FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies
by: Hu, Xintong, et al.
Published: (2026)
by: Hu, Xintong, et al.
Published: (2026)
Commonsense Reasoning for Legged Robot Adaptation with Vision-Language Models
by: Chen, Annie S., et al.
Published: (2024)
by: Chen, Annie S., et al.
Published: (2024)
DexHub and DART: Towards Internet Scale Robot Data Collection
by: Park, Younghyo, et al.
Published: (2024)
by: Park, Younghyo, et al.
Published: (2024)
$π_0$: A Vision-Language-Action Flow Model for General Robot Control
by: Black, Kevin, et al.
Published: (2024)
by: Black, Kevin, et al.
Published: (2024)
La sécurité du travail est aussi une affaire de comportement
by: Jagdeep Singh Chhokar
Published: (1987)
by: Jagdeep Singh Chhokar
Published: (1987)
Real-Time Execution of Action Chunking Flow Policies
by: Black, Kevin, et al.
Published: (2025)
by: Black, Kevin, et al.
Published: (2025)
Time for a more sustainable and hopeful vision of bovine TB freedom for the UK
by: Christianne Glossop
Published: (2025)
by: Christianne Glossop
Published: (2025)
Pushing the Limits of Cross-Embodiment Learning for Manipulation and Navigation
by: Yang, Jonathan, et al.
Published: (2024)
by: Yang, Jonathan, et al.
Published: (2024)
Unmasking the Illusion of Embodied Reasoning in Vision-Language-Action Models
by: Xu, Haiweng, et al.
Published: (2026)
by: Xu, Haiweng, et al.
Published: (2026)
Vlaser: Vision-Language-Action Model with Synergistic Embodied Reasoning
by: Yang, Ganlin, et al.
Published: (2025)
by: Yang, Ganlin, et al.
Published: (2025)
Structural Bounds and Forbidden Induced Subgraphs for Edge-Add Graph Classes
by: Singh, Jagdeep, et al.
Published: (2025)
by: Singh, Jagdeep, et al.
Published: (2025)
Loose elements in binary and ternary matroids
by: Singh, Jagdeep, et al.
Published: (2025)
by: Singh, Jagdeep, et al.
Published: (2025)
Extensions and Deletions of matroid classes closed under flats
by: Singh, Jagdeep, et al.
Published: (2024)
by: Singh, Jagdeep, et al.
Published: (2024)
Edge-apexing in hereditary classes of graphs
by: Singh, Jagdeep, et al.
Published: (2024)
by: Singh, Jagdeep, et al.
Published: (2024)
Learning Affordances at Inference-Time for Vision-Language-Action Models
by: Shah, Ameesh, et al.
Published: (2025)
by: Shah, Ameesh, et al.
Published: (2025)
SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities
by: Chen, Boyuan, et al.
Published: (2024)
by: Chen, Boyuan, et al.
Published: (2024)
OpenVLA: An Open-Source Vision-Language-Action Model
by: Kim, Moo Jin, et al.
Published: (2024)
by: Kim, Moo Jin, et al.
Published: (2024)
Optimization‐Driven Localization in Wireless Sensor Networks: A Comprehensive Review of Single and Hybrid Metaheuristic Approaches
by: Tajinder Kaur, et al.
Published: (2025)
by: Tajinder Kaur, et al.
Published: (2025)
Guide, Think, Act: Interactive Embodied Reasoning in Vision-Language-Action Models
by: Ling, Yiran, et al.
Published: (2026)
by: Ling, Yiran, et al.
Published: (2026)
Embodied-Reasoner: Synergizing Visual Search, Reasoning, and Action for Embodied Interactive Tasks
by: Zhang, Wenqi, et al.
Published: (2025)
by: Zhang, Wenqi, et al.
Published: (2025)
Similar Items
-
Training Strategies for Efficient Embodied Reasoning
by: Chen, William, et al.
Published: (2025) -
FAST: Efficient Action Tokenization for Vision-Language-Action Models
by: Pertsch, Karl, et al.
Published: (2025) -
OmniVLA: An Omni-Modal Vision-Language-Action Model for Robot Navigation
by: Hirose, Noriaki, et al.
Published: (2025) -
Scaling Cross-Embodied Learning: One Policy for Manipulation, Navigation, Locomotion and Aviation
by: Doshi, Ria, et al.
Published: (2024) -
CAST: Counterfactual Labels Improve Instruction Following in Vision-Language-Action Models
by: Glossop, Catherine, et al.
Published: (2025)