Think, Remember, Navigate: Zero-Shot Object-Goal Navigation with VLM-Powered Reasoning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Habibpour, Mobin, Afghah, Fatemeh
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914153021571072
author Habibpour, Mobin
Afghah, Fatemeh
author_facet Habibpour, Mobin
Afghah, Fatemeh
contents While Vision-Language Models (VLMs) are set to transform robotic navigation, existing methods often underutilize their reasoning capabilities. To unlock the full potential of VLMs in robotics, we shift their role from passive observers to active strategists in the navigation process. Our framework outsources high-level planning to a VLM, which leverages its contextual understanding to guide a frontier-based exploration agent. This intelligent guidance is achieved through a trio of techniques: structured chain-of-thought prompting that elicits logical, step-by-step reasoning; dynamic inclusion of the agent's recent action history to prevent getting stuck in loops; and a novel capability that enables the VLM to interpret top-down obstacle maps alongside first-person views, thereby enhancing spatial awareness. When tested on challenging benchmarks like HM3D, Gibson, and MP3D, this method produces exceptionally direct and logical trajectories, marking a substantial improvement in navigation efficiency over existing approaches and charting a path toward more capable embodied agents.
format Preprint
id arxiv_https___arxiv_org_abs_2511_08942
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Think, Remember, Navigate: Zero-Shot Object-Goal Navigation with VLM-Powered Reasoning
Habibpour, Mobin
Afghah, Fatemeh
Robotics
Artificial Intelligence
While Vision-Language Models (VLMs) are set to transform robotic navigation, existing methods often underutilize their reasoning capabilities. To unlock the full potential of VLMs in robotics, we shift their role from passive observers to active strategists in the navigation process. Our framework outsources high-level planning to a VLM, which leverages its contextual understanding to guide a frontier-based exploration agent. This intelligent guidance is achieved through a trio of techniques: structured chain-of-thought prompting that elicits logical, step-by-step reasoning; dynamic inclusion of the agent's recent action history to prevent getting stuck in loops; and a novel capability that enables the VLM to interpret top-down obstacle maps alongside first-person views, thereby enhancing spatial awareness. When tested on challenging benchmarks like HM3D, Gibson, and MP3D, this method produces exceptionally direct and logical trajectories, marking a substantial improvement in navigation efficiency over existing approaches and charting a path toward more capable embodied agents.
title Think, Remember, Navigate: Zero-Shot Object-Goal Navigation with VLM-Powered Reasoning
topic Robotics
Artificial Intelligence
url https://arxiv.org/abs/2511.08942