NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Cheng, An-Chieh, Ji, Yandong, Yang, Zhaojing, Gongye, Zaitian, Zou, Xueyan, Kautz, Jan, Bıyık, Erdem, Yin, Hongxu, Liu, Sifei, Wang, Xiaolong
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909497112395776
author Cheng, An-Chieh
Ji, Yandong
Yang, Zhaojing
Gongye, Zaitian
Zou, Xueyan
Kautz, Jan
Bıyık, Erdem
Yin, Hongxu
Liu, Sifei
Wang, Xiaolong
author_facet Cheng, An-Chieh
Ji, Yandong
Yang, Zhaojing
Gongye, Zaitian
Zou, Xueyan
Kautz, Jan
Bıyık, Erdem
Yin, Hongxu
Liu, Sifei
Wang, Xiaolong
contents This paper proposes to solve the problem of Vision-and-Language Navigation with legged robots, which not only provides a flexible way for humans to command but also allows the robot to navigate through more challenging and cluttered scenes. However, it is non-trivial to translate human language instructions all the way to low-level leg joint actions. We propose NaVILA, a 2-level framework that unifies a Vision-Language-Action model (VLA) with locomotion skills. Instead of directly predicting low-level actions from VLA, NaVILA first generates mid-level actions with spatial information in the form of language, (e.g., "moving forward 75cm"), which serves as an input for a visual locomotion RL policy for execution. NaVILA substantially improves previous approaches on existing benchmarks. The same advantages are demonstrated in our newly developed benchmarks with IsaacLab, featuring more realistic scenes, low-level controls, and real-world robot experiments. We show more results at https://navila-bot.github.io/
format Preprint
id arxiv_https___arxiv_org_abs_2412_04453
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle NaVILA: Legged Robot Vision-Language-Action Model for Navigation
Cheng, An-Chieh
Ji, Yandong
Yang, Zhaojing
Gongye, Zaitian
Zou, Xueyan
Kautz, Jan
Bıyık, Erdem
Yin, Hongxu
Liu, Sifei
Wang, Xiaolong
Robotics
Computer Vision and Pattern Recognition
This paper proposes to solve the problem of Vision-and-Language Navigation with legged robots, which not only provides a flexible way for humans to command but also allows the robot to navigate through more challenging and cluttered scenes. However, it is non-trivial to translate human language instructions all the way to low-level leg joint actions. We propose NaVILA, a 2-level framework that unifies a Vision-Language-Action model (VLA) with locomotion skills. Instead of directly predicting low-level actions from VLA, NaVILA first generates mid-level actions with spatial information in the form of language, (e.g., "moving forward 75cm"), which serves as an input for a visual locomotion RL policy for execution. NaVILA substantially improves previous approaches on existing benchmarks. The same advantages are demonstrated in our newly developed benchmarks with IsaacLab, featuring more realistic scenes, low-level controls, and real-world robot experiments. We show more results at https://navila-bot.github.io/
title NaVILA: Legged Robot Vision-Language-Action Model for Navigation
topic Robotics
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.04453