Zero-shot Object Navigation with Vision-Language Models Reasoning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wen, Congcong, Huang, Yisiyuan, Huang, Hao, Huang, Yanjia, Yuan, Shuaihang, Hao, Yu, Lin, Hui, Liu, Yu-Shen, Fang, Yi
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913561796673536
author Wen, Congcong
Huang, Yisiyuan
Huang, Hao
Huang, Yanjia
Yuan, Shuaihang
Hao, Yu
Lin, Hui
Liu, Yu-Shen
Fang, Yi
author_facet Wen, Congcong
Huang, Yisiyuan
Huang, Hao
Huang, Yanjia
Yuan, Shuaihang
Hao, Yu
Lin, Hui
Liu, Yu-Shen
Fang, Yi
contents Object navigation is crucial for robots, but traditional methods require substantial training data and cannot be generalized to unknown environments. Zero-shot object navigation (ZSON) aims to address this challenge, allowing robots to interact with unknown objects without specific training data. Language-driven zero-shot object navigation (L-ZSON) is an extension of ZSON that incorporates natural language instructions to guide robot navigation and interaction with objects. In this paper, we propose a novel Vision Language model with a Tree-of-thought Network (VLTNet) for L-ZSON. VLTNet comprises four main modules: vision language model understanding, semantic mapping, tree-of-thought reasoning and exploration, and goal identification. Among these modules, Tree-of-Thought (ToT) reasoning and exploration module serves as a core component, innovatively using the ToT reasoning framework for navigation frontier selection during robot exploration. Compared to conventional frontier selection without reasoning, navigation using ToT reasoning involves multi-path reasoning processes and backtracking when necessary, enabling globally informed decision-making with higher accuracy. Experimental results on PASTURE and RoboTHOR benchmarks demonstrate the outstanding performance of our model in LZSON, particularly in scenarios involving complex natural language as target instructions.
format Preprint
id arxiv_https___arxiv_org_abs_2410_18570
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Zero-shot Object Navigation with Vision-Language Models Reasoning
Wen, Congcong
Huang, Yisiyuan
Huang, Hao
Huang, Yanjia
Yuan, Shuaihang
Hao, Yu
Lin, Hui
Liu, Yu-Shen
Fang, Yi
Robotics
Artificial Intelligence
Object navigation is crucial for robots, but traditional methods require substantial training data and cannot be generalized to unknown environments. Zero-shot object navigation (ZSON) aims to address this challenge, allowing robots to interact with unknown objects without specific training data. Language-driven zero-shot object navigation (L-ZSON) is an extension of ZSON that incorporates natural language instructions to guide robot navigation and interaction with objects. In this paper, we propose a novel Vision Language model with a Tree-of-thought Network (VLTNet) for L-ZSON. VLTNet comprises four main modules: vision language model understanding, semantic mapping, tree-of-thought reasoning and exploration, and goal identification. Among these modules, Tree-of-Thought (ToT) reasoning and exploration module serves as a core component, innovatively using the ToT reasoning framework for navigation frontier selection during robot exploration. Compared to conventional frontier selection without reasoning, navigation using ToT reasoning involves multi-path reasoning processes and backtracking when necessary, enabling globally informed decision-making with higher accuracy. Experimental results on PASTURE and RoboTHOR benchmarks demonstrate the outstanding performance of our model in LZSON, particularly in scenarios involving complex natural language as target instructions.
title Zero-shot Object Navigation with Vision-Language Models Reasoning
topic Robotics
Artificial Intelligence
url https://arxiv.org/abs/2410.18570