HiRT: Enhancing Robotic Control with Hierarchical Robot Transformers
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866929694524309504 |
|---|---|
| author | Zhang, Jianke Guo, Yanjiang Chen, Xiaoyu Wang, Yen-Jen Hu, Yucheng Shi, Chengming Chen, Jianyu |
| author_facet | Zhang, Jianke Guo, Yanjiang Chen, Xiaoyu Wang, Yen-Jen Hu, Yucheng Shi, Chengming Chen, Jianyu |
| contents | Large Vision-Language-Action (VLA) models, leveraging powerful pre trained Vision-Language Models (VLMs) backends, have shown promise in robotic control due to their impressive generalization ability. However, the success comes at a cost. Their reliance on VLM backends with billions of parameters leads to high computational costs and inference latency, limiting the testing scenarios to mainly quasi-static tasks and hindering performance in dynamic tasks requiring rapid interactions. To address these limitations, this paper proposes HiRT, a Hierarchical Robot Transformer framework that enables flexible frequency and performance trade-off. HiRT keeps VLMs running at low frequencies to capture temporarily invariant features while enabling real-time interaction through a high-frequency vision-based policy guided by the slowly updated features. Experiment results in both simulation and real-world settings demonstrate significant improvements over baseline methods. Empirically, in static tasks, we double the control frequency and achieve comparable success rates. Additionally, on novel real-world dynamic ma nipulation tasks which are challenging for previous VLA models, HiRT improves the success rate from 48% to 75%. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2410_05273 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | HiRT: Enhancing Robotic Control with Hierarchical Robot Transformers Zhang, Jianke Guo, Yanjiang Chen, Xiaoyu Wang, Yen-Jen Hu, Yucheng Shi, Chengming Chen, Jianyu Computer Vision and Pattern Recognition Artificial Intelligence Robotics Large Vision-Language-Action (VLA) models, leveraging powerful pre trained Vision-Language Models (VLMs) backends, have shown promise in robotic control due to their impressive generalization ability. However, the success comes at a cost. Their reliance on VLM backends with billions of parameters leads to high computational costs and inference latency, limiting the testing scenarios to mainly quasi-static tasks and hindering performance in dynamic tasks requiring rapid interactions. To address these limitations, this paper proposes HiRT, a Hierarchical Robot Transformer framework that enables flexible frequency and performance trade-off. HiRT keeps VLMs running at low frequencies to capture temporarily invariant features while enabling real-time interaction through a high-frequency vision-based policy guided by the slowly updated features. Experiment results in both simulation and real-world settings demonstrate significant improvements over baseline methods. Empirically, in static tasks, we double the control frequency and achieve comparable success rates. Additionally, on novel real-world dynamic ma nipulation tasks which are challenging for previous VLA models, HiRT improves the success rate from 48% to 75%. |
| title | HiRT: Enhancing Robotic Control with Hierarchical Robot Transformers |
| topic | Computer Vision and Pattern Recognition Artificial Intelligence Robotics |
| url | https://arxiv.org/abs/2410.05273 |