NavGPT-2: Unleashing Navigational Reasoning Capability for Large Vision-Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhou, Gengze, Hong, Yicong, Wang, Zun, Wang, Xin Eric, Wu, Qi
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910613398093824
author Zhou, Gengze
Hong, Yicong
Wang, Zun
Wang, Xin Eric
Wu, Qi
author_facet Zhou, Gengze
Hong, Yicong
Wang, Zun
Wang, Xin Eric
Wu, Qi
contents Capitalizing on the remarkable advancements in Large Language Models (LLMs), there is a burgeoning initiative to harness LLMs for instruction following robotic navigation. Such a trend underscores the potential of LLMs to generalize navigational reasoning and diverse language understanding. However, a significant discrepancy in agent performance is observed when integrating LLMs in the Vision-and-Language navigation (VLN) tasks compared to previous downstream specialist models. Furthermore, the inherent capacity of language to interpret and facilitate communication in agent interactions is often underutilized in these integrations. In this work, we strive to bridge the divide between VLN-specialized models and LLM-based navigation paradigms, while maintaining the interpretative prowess of LLMs in generating linguistic navigational reasoning. By aligning visual content in a frozen LLM, we encompass visual observation comprehension for LLMs and exploit a way to incorporate LLMs and navigation policy networks for effective action predictions and navigational reasoning. We demonstrate the data efficiency of the proposed methods and eliminate the gap between LM-based agents and state-of-the-art VLN specialists.
format Preprint
id arxiv_https___arxiv_org_abs_2407_12366
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle NavGPT-2: Unleashing Navigational Reasoning Capability for Large Vision-Language Models
Zhou, Gengze
Hong, Yicong
Wang, Zun
Wang, Xin Eric
Wu, Qi
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Robotics
Capitalizing on the remarkable advancements in Large Language Models (LLMs), there is a burgeoning initiative to harness LLMs for instruction following robotic navigation. Such a trend underscores the potential of LLMs to generalize navigational reasoning and diverse language understanding. However, a significant discrepancy in agent performance is observed when integrating LLMs in the Vision-and-Language navigation (VLN) tasks compared to previous downstream specialist models. Furthermore, the inherent capacity of language to interpret and facilitate communication in agent interactions is often underutilized in these integrations. In this work, we strive to bridge the divide between VLN-specialized models and LLM-based navigation paradigms, while maintaining the interpretative prowess of LLMs in generating linguistic navigational reasoning. By aligning visual content in a frozen LLM, we encompass visual observation comprehension for LLMs and exploit a way to incorporate LLMs and navigation policy networks for effective action predictions and navigational reasoning. We demonstrate the data efficiency of the proposed methods and eliminate the gap between LM-based agents and state-of-the-art VLN specialists.
title NavGPT-2: Unleashing Navigational Reasoning Capability for Large Vision-Language Models
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Robotics
url https://arxiv.org/abs/2407.12366