HiRT: Enhancing Robotic Control with Hierarchical Robot Transformers

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Zhang, Jianke, Guo, Yanjiang, Chen, Xiaoyu, Wang, Yen-Jen, Hu, Yucheng, Shi, Chengming, Chen, Jianyu
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866929694524309504
author Zhang, Jianke
Guo, Yanjiang
Chen, Xiaoyu
Wang, Yen-Jen
Hu, Yucheng
Shi, Chengming
Chen, Jianyu
author_facet Zhang, Jianke
Guo, Yanjiang
Chen, Xiaoyu
Wang, Yen-Jen
Hu, Yucheng
Shi, Chengming
Chen, Jianyu
contents Large Vision-Language-Action (VLA) models, leveraging powerful pre trained Vision-Language Models (VLMs) backends, have shown promise in robotic control due to their impressive generalization ability. However, the success comes at a cost. Their reliance on VLM backends with billions of parameters leads to high computational costs and inference latency, limiting the testing scenarios to mainly quasi-static tasks and hindering performance in dynamic tasks requiring rapid interactions. To address these limitations, this paper proposes HiRT, a Hierarchical Robot Transformer framework that enables flexible frequency and performance trade-off. HiRT keeps VLMs running at low frequencies to capture temporarily invariant features while enabling real-time interaction through a high-frequency vision-based policy guided by the slowly updated features. Experiment results in both simulation and real-world settings demonstrate significant improvements over baseline methods. Empirically, in static tasks, we double the control frequency and achieve comparable success rates. Additionally, on novel real-world dynamic ma nipulation tasks which are challenging for previous VLA models, HiRT improves the success rate from 48% to 75%.
format Preprint
id arxiv_https___arxiv_org_abs_2410_05273
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle HiRT: Enhancing Robotic Control with Hierarchical Robot Transformers
Zhang, Jianke
Guo, Yanjiang
Chen, Xiaoyu
Wang, Yen-Jen
Hu, Yucheng
Shi, Chengming
Chen, Jianyu
Computer Vision and Pattern Recognition
Artificial Intelligence
Robotics
Large Vision-Language-Action (VLA) models, leveraging powerful pre trained Vision-Language Models (VLMs) backends, have shown promise in robotic control due to their impressive generalization ability. However, the success comes at a cost. Their reliance on VLM backends with billions of parameters leads to high computational costs and inference latency, limiting the testing scenarios to mainly quasi-static tasks and hindering performance in dynamic tasks requiring rapid interactions. To address these limitations, this paper proposes HiRT, a Hierarchical Robot Transformer framework that enables flexible frequency and performance trade-off. HiRT keeps VLMs running at low frequencies to capture temporarily invariant features while enabling real-time interaction through a high-frequency vision-based policy guided by the slowly updated features. Experiment results in both simulation and real-world settings demonstrate significant improvements over baseline methods. Empirically, in static tasks, we double the control frequency and achieve comparable success rates. Additionally, on novel real-world dynamic ma nipulation tasks which are challenging for previous VLA models, HiRT improves the success rate from 48% to 75%.
title HiRT: Enhancing Robotic Control with Hierarchical Robot Transformers
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Robotics
url https://arxiv.org/abs/2410.05273