RT-H: Action Hierarchies Using Language
Fuente:
arXiv
Saved in:
| Main Authors: | Belkhale, Suneel, Ding, Tianli, Xiao, Ted, Sermanet, Pierre, Vuong, Quon, Tompson, Jonathan, Chebotar, Yevgen, Dwibedi, Debidatta, Sadigh, Dorsa |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Action-Free Reasoning for Policy Generalization
by: Clark, Jaden, et al.
Published: (2025)
by: Clark, Jaden, et al.
Published: (2025)
A Short Note on Evaluating RepNet for Temporal Repetition Counting in Videos
by: Dwibedi, Debidatta, et al.
Published: (2024)
by: Dwibedi, Debidatta, et al.
Published: (2024)
A Taxonomy for Evaluating Generalist Robot Manipulation Policies
by: Gao, Jensen, et al.
Published: (2025)
by: Gao, Jensen, et al.
Published: (2025)
So You Think You Can Scale Up Autonomous Robot Data Collection?
by: Mirchandani, Suvir, et al.
Published: (2024)
by: Mirchandani, Suvir, et al.
Published: (2024)
RT-Affordance: Affordances are Versatile Intermediate Representations for Robot Manipulation
by: Nasiriany, Soroush, et al.
Published: (2024)
by: Nasiriany, Soroush, et al.
Published: (2024)
Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation
by: Bharadhwaj, Homanga, et al.
Published: (2024)
by: Bharadhwaj, Homanga, et al.
Published: (2024)
RT-Sketch: Goal-Conditioned Imitation Learning from Hand-Drawn Sketches
by: Sundaresan, Priya, et al.
Published: (2024)
by: Sundaresan, Priya, et al.
Published: (2024)
AutoRT: Embodied Foundation Models for Large Scale Orchestration of Robotic Agents
by: Ahn, Michael, et al.
Published: (2024)
by: Ahn, Michael, et al.
Published: (2024)
OVR: A Dataset for Open Vocabulary Temporal Repetition Counting in Videos
by: Dwibedi, Debidatta, et al.
Published: (2024)
by: Dwibedi, Debidatta, et al.
Published: (2024)
Efficient Data Collection for Robotic Manipulation via Compositional Generalization
by: Gao, Jensen, et al.
Published: (2024)
by: Gao, Jensen, et al.
Published: (2024)
Vid2Robot: End-to-end Video-conditioned Policy Learning with Cross-Attention Transformers
by: Jain, Vidhi, et al.
Published: (2024)
by: Jain, Vidhi, et al.
Published: (2024)
Unified Video Action Model
by: Li, Shuang, et al.
Published: (2025)
by: Li, Shuang, et al.
Published: (2025)
Efficiently Generating Expressive Quadruped Behaviors via Language-Guided Preference Learning
by: Clark, Jaden, et al.
Published: (2025)
by: Clark, Jaden, et al.
Published: (2025)
FlexCap: Describe Anything in Images in Controllable Detail
by: Dwibedi, Debidatta, et al.
Published: (2024)
by: Dwibedi, Debidatta, et al.
Published: (2024)
Data Analogies Enable Efficient Cross-Embodiment Transfer
by: Yang, Jonathan, et al.
Published: (2026)
by: Yang, Jonathan, et al.
Published: (2026)
EXPO-FT: Sample-Efficient Reinforcement Learning Finetuning for Vision-Language-Action Models
by: Dong, Perry, et al.
Published: (2026)
by: Dong, Perry, et al.
Published: (2026)
Invariance Co-training for Robot Visual Generalization
by: Yang, Jonathan, et al.
Published: (2025)
by: Yang, Jonathan, et al.
Published: (2025)
Robot Data Curation with Mutual Information Estimators
by: Hejna, Joey, et al.
Published: (2025)
by: Hejna, Joey, et al.
Published: (2025)
HandelBot: Real-World Piano Playing via Fast Adaptation of Dexterous Robot Policies
by: Xie, Amber, et al.
Published: (2026)
by: Xie, Amber, et al.
Published: (2026)
Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization
by: Yang, Jonathan, et al.
Published: (2025)
by: Yang, Jonathan, et al.
Published: (2025)
What's the Move? Hybrid Imitation Learning via Salient Points
by: Sundaresan, Priya, et al.
Published: (2024)
by: Sundaresan, Priya, et al.
Published: (2024)
OpenVLA: An Open-Source Vision-Language-Action Model
by: Kim, Moo Jin, et al.
Published: (2024)
by: Kim, Moo Jin, et al.
Published: (2024)
How to Train Your Robots? The Impact of Demonstration Modality on Imitation Learning
by: Li, Haozhuo, et al.
Published: (2025)
by: Li, Haozhuo, et al.
Published: (2025)
Physically Grounded Vision-Language Models for Robotic Manipulation
by: Gao, Jensen, et al.
Published: (2023)
by: Gao, Jensen, et al.
Published: (2023)
Training Strategies for Efficient Embodied Reasoning
by: Chen, William, et al.
Published: (2025)
by: Chen, William, et al.
Published: (2025)
Batch Active Learning of Reward Functions from Human Preferences
by: Bıyık, Erdem, et al.
Published: (2024)
by: Bıyık, Erdem, et al.
Published: (2024)
Grounding Robot Generalization in Training Data via Retrieval-Augmented VLMs
by: Gao, Jensen, et al.
Published: (2026)
by: Gao, Jensen, et al.
Published: (2026)
Will People Enjoy a Robot Trainer? A Case Study with Snoopie the Pacerbot
by: Du, Maximilian, et al.
Published: (2026)
by: Du, Maximilian, et al.
Published: (2026)
Pushing the Limits of Cross-Embodiment Learning for Manipulation and Navigation
by: Yang, Jonathan, et al.
Published: (2024)
by: Yang, Jonathan, et al.
Published: (2024)
Vision Language Models are In-Context Value Learners
by: Ma, Yecheng Jason, et al.
Published: (2024)
by: Ma, Yecheng Jason, et al.
Published: (2024)
Latent Diffusion Planning for Imitation Learning
by: Xie, Amber, et al.
Published: (2025)
by: Xie, Amber, et al.
Published: (2025)
Data Retrieval with Importance Weights for Few-Shot Imitation Learning
by: Xie, Amber, et al.
Published: (2025)
by: Xie, Amber, et al.
Published: (2025)
What Matters for Batch Online Reinforcement Learning in Robotics?
by: Dong, Perry, et al.
Published: (2025)
by: Dong, Perry, et al.
Published: (2025)
BPP: Long-Context Robot Imitation Learning by Focusing on Key History Frames
by: Mark, Max Sobol, et al.
Published: (2026)
by: Mark, Max Sobol, et al.
Published: (2026)
Altruistic Maneuver Planning for Cooperative Autonomous Vehicles Using Multi-agent Advantage Actor-Critic
by: Toghi, Behrad, et al.
Published: (2021)
by: Toghi, Behrad, et al.
Published: (2021)
SteerVLA: Steering Vision-Language-Action Models in Long-Tail Driving Scenarios
by: Gao, Tian, et al.
Published: (2026)
by: Gao, Tian, et al.
Published: (2026)
Generative Expressive Robot Behaviors using Large Language Models
by: Mahadevan, Karthik, et al.
Published: (2024)
by: Mahadevan, Karthik, et al.
Published: (2024)
DexDrummer: In-Hand, Contact-Rich, and Long-Horizon Dexterous Robot Drumming
by: Fang, Hung-Chieh, et al.
Published: (2026)
by: Fang, Hung-Chieh, et al.
Published: (2026)
Re-Mix: Optimizing Data Mixtures for Large Scale Imitation Learning
by: Hejna, Joey, et al.
Published: (2024)
by: Hejna, Joey, et al.
Published: (2024)
MotIF: Motion Instruction Fine-tuning
by: Hwang, Minyoung, et al.
Published: (2024)
by: Hwang, Minyoung, et al.
Published: (2024)
Similar Items
-
Action-Free Reasoning for Policy Generalization
by: Clark, Jaden, et al.
Published: (2025) -
A Short Note on Evaluating RepNet for Temporal Repetition Counting in Videos
by: Dwibedi, Debidatta, et al.
Published: (2024) -
A Taxonomy for Evaluating Generalist Robot Manipulation Policies
by: Gao, Jensen, et al.
Published: (2025) -
So You Think You Can Scale Up Autonomous Robot Data Collection?
by: Mirchandani, Suvir, et al.
Published: (2024) -
RT-Affordance: Affordances are Versatile Intermediate Representations for Robot Manipulation
by: Nasiriany, Soroush, et al.
Published: (2024)