Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models
Fuente:
arXiv
Saved in:
| Main Authors: | Shi, Lucy Xiaoyang, Ichter, Brian, Equi, Michael, Ke, Liyiming, Pertsch, Karl, Vuong, Quan, Tanner, James, Walling, Anna, Wang, Haohuan, Fusai, Niccolo, Li-Bell, Adrian, Driess, Danny, Groom, Lachy, Levine, Sergey, Finn, Chelsea |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
$π_0$: A Vision-Language-Action Flow Model for General Robot Control
by: Black, Kevin, et al.
Published: (2024)
by: Black, Kevin, et al.
Published: (2024)
FAST: Efficient Action Tokenization for Vision-Language-Action Models
by: Pertsch, Karl, et al.
Published: (2025)
by: Pertsch, Karl, et al.
Published: (2025)
$π_{0.5}$: a Vision-Language-Action Model with Open-World Generalization
by: Intelligence, Physical, et al.
Published: (2025)
by: Intelligence, Physical, et al.
Published: (2025)
Knowledge Insulating Vision-Language-Action Models: Train Fast, Run Fast, Generalize Better
by: Driess, Danny, et al.
Published: (2025)
by: Driess, Danny, et al.
Published: (2025)
MEM: Multi-Scale Embodied Memory for Vision Language Action Models
by: Torne, Marcel, et al.
Published: (2026)
by: Torne, Marcel, et al.
Published: (2026)
Yell At Your Robot: Improving On-the-Fly from Language Corrections
by: Shi, Lucy Xiaoyang, et al.
Published: (2024)
by: Shi, Lucy Xiaoyang, et al.
Published: (2024)
Training Strategies for Efficient Embodied Reasoning
by: Chen, William, et al.
Published: (2025)
by: Chen, William, et al.
Published: (2025)
$π^{*}_{0.6}$: a VLA That Learns From Experience
by: Intelligence, Physical, et al.
Published: (2025)
by: Intelligence, Physical, et al.
Published: (2025)
RoboReward: General-Purpose Vision-Language Reward Models for Robotics
by: Lee, Tony, et al.
Published: (2026)
by: Lee, Tony, et al.
Published: (2026)
Robotic Control via Embodied Chain-of-Thought Reasoning
by: Zawalski, Michał, et al.
Published: (2024)
by: Zawalski, Michał, et al.
Published: (2024)
Long-Horizon Q-Learning: Accurate Value Learning via n-Step Inequalities
by: Abraham, Armaan A., et al.
Published: (2026)
by: Abraham, Armaan A., et al.
Published: (2026)
RL Token: Bootstrapping Online RL with Vision-Language-Action Models
by: Xu, Charles, et al.
Published: (2026)
by: Xu, Charles, et al.
Published: (2026)
Feedback Descent: Open-Ended Text Optimization via Pairwise Comparison
by: Lee, Yoonho, et al.
Published: (2025)
by: Lee, Yoonho, et al.
Published: (2025)
Ctrl-World: A Controllable Generative World Model for Robot Manipulation
by: Guo, Yanjiang, et al.
Published: (2025)
by: Guo, Yanjiang, et al.
Published: (2025)
Pushing the Limits of Cross-Embodiment Learning for Manipulation and Navigation
by: Yang, Jonathan, et al.
Published: (2024)
by: Yang, Jonathan, et al.
Published: (2024)
ALOHA Unleashed: A Simple Recipe for Robot Dexterity
by: Zhao, Tony Z., et al.
Published: (2024)
by: Zhao, Tony Z., et al.
Published: (2024)
Steerable Vision-Language-Action Policies for Embodied Reasoning and Hierarchical Control
by: Chen, William, et al.
Published: (2026)
by: Chen, William, et al.
Published: (2026)
Emergence of Human to Robot Transfer in Vision-Language-Action Models
by: Kareer, Simar, et al.
Published: (2025)
by: Kareer, Simar, et al.
Published: (2025)
Affordance-Guided Reinforcement Learning via Visual Prompting
by: Lee, Olivia Y., et al.
Published: (2024)
by: Lee, Olivia Y., et al.
Published: (2024)
PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs
by: Nasiriany, Soroush, et al.
Published: (2024)
by: Nasiriany, Soroush, et al.
Published: (2024)
SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities
by: Chen, Boyuan, et al.
Published: (2024)
by: Chen, Boyuan, et al.
Published: (2024)
VLAW: Iterative Co-Improvement of Vision-Language-Action Policy and World Model
by: Guo, Yanjiang, et al.
Published: (2026)
by: Guo, Yanjiang, et al.
Published: (2026)
Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models
by: Zhang, Zhicheng, et al.
Published: (2025)
by: Zhang, Zhicheng, et al.
Published: (2025)
Training-Time Action Conditioning for Efficient Real-Time Chunking
by: Black, Kevin, et al.
Published: (2025)
by: Black, Kevin, et al.
Published: (2025)
Evaluating Real-World Robot Manipulation Policies in Simulation
by: Li, Xuanlin, et al.
Published: (2024)
by: Li, Xuanlin, et al.
Published: (2024)
Posterior Behavioral Cloning: Pretraining BC Policies for Efficient RL Finetuning
by: Wagenmaker, Andrew, et al.
Published: (2025)
by: Wagenmaker, Andrew, et al.
Published: (2025)
CrafText Benchmark: Advancing Instruction Following in Complex Multimodal Open-Ended World
by: Volovikova, Zoya, et al.
Published: (2025)
by: Volovikova, Zoya, et al.
Published: (2025)
[Fresno County Library Rural Literacy Outreach Program. Final Performance Report, 1988-1989.]
by: Walling, Joyce
Published: (1990)
by: Walling, Joyce
Published: (1990)
CHEMICAL CONTROL OVER THE SYNTHESIS OF DEFECT-FREE AND HETEROMETAL-DOPED METAL OXIDE NANOPARTICLES BY THE SINGLE-SOURCE PRECURSOR APPROACH
by: M. Driess
Published: (2005)
by: M. Driess
Published: (2005)
OpenVLA: An Open-Source Vision-Language-Action Model
by: Kim, Moo Jin, et al.
Published: (2024)
by: Kim, Moo Jin, et al.
Published: (2024)
$π_{0.7}$: a Steerable Generalist Robotic Foundation Model with Emergent Capabilities
by: Intelligence, Physical, et al.
Published: (2026)
by: Intelligence, Physical, et al.
Published: (2026)
Improving claims handling with predictive analytics / Robert J. Walling
by: Walling, Robert J
by: Walling, Robert J
Ability, Disability, and Picture Books.
by: Walling, Linda Lucas
Published: (2001)
by: Walling, Linda Lucas
Published: (2001)
Going the Distance: Equal Education, Off Campus or On.
by: Walling, Linda Lucas
Published: (1996)
by: Walling, Linda Lucas
Published: (1996)
Public Libraries and People with Mental Retardation.
by: Walling, Linda Lucas
Published: (2001)
by: Walling, Linda Lucas
Published: (2001)
Granting Each Equal Access.
by: Walling, Linda Lucas
Published: (1992)
by: Walling, Linda Lucas
Published: (1992)
Montego Bay’s Marine Park: the real bottom line
by: Walling, Leslie J.
Published: (1994)
by: Walling, Leslie J.
Published: (1994)
Alternative Price Dynamics and Valuation of Flexible Strategies
by: Cristina Bertolosi, et al.
Published: (2025)
by: Cristina Bertolosi, et al.
Published: (2025)
Autonomous Improvement of Instruction Following Skills via Foundation Models
by: Zhou, Zhiyuan, et al.
Published: (2024)
by: Zhou, Zhiyuan, et al.
Published: (2024)
Octo: An Open-Source Generalist Robot Policy
by: Octo Model Team, et al.
Published: (2024)
by: Octo Model Team, et al.
Published: (2024)
Similar Items
-
$π_0$: A Vision-Language-Action Flow Model for General Robot Control
by: Black, Kevin, et al.
Published: (2024) -
FAST: Efficient Action Tokenization for Vision-Language-Action Models
by: Pertsch, Karl, et al.
Published: (2025) -
$π_{0.5}$: a Vision-Language-Action Model with Open-World Generalization
by: Intelligence, Physical, et al.
Published: (2025) -
Knowledge Insulating Vision-Language-Action Models: Train Fast, Run Fast, Generalize Better
by: Driess, Danny, et al.
Published: (2025) -
MEM: Multi-Scale Embodied Memory for Vision Language Action Models
by: Torne, Marcel, et al.
Published: (2026)