MAG-Nav: Language-Driven Object Navigation Leveraging Memory-Reserved Active Grounding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Weifan, Li, Tingguang, Liu, Yuzhen
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918116495196160
author Zhang, Weifan
Li, Tingguang
Liu, Yuzhen
author_facet Zhang, Weifan
Li, Tingguang
Liu, Yuzhen
contents Visual navigation in unknown environments based solely on natural language descriptions is a key capability for intelligent robots. In this work, we propose a navigation framework built upon off-the-shelf Visual Language Models (VLMs), enhanced with two human-inspired mechanisms: perspective-based active grounding, which dynamically adjusts the robot's viewpoint for improved visual inspection, and historical memory backtracking, which enables the system to retain and re-evaluate uncertain observations over time. Unlike existing approaches that passively rely on incidental visual inputs, our method actively optimizes perception and leverages memory to resolve ambiguity, significantly improving vision-language grounding in complex, unseen environments. Our framework operates in a zero-shot manner, achieving strong generalization to diverse and open-ended language descriptions without requiring labeled data or model fine-tuning. Experimental results on Habitat-Matterport 3D (HM3D) show that our method outperforms state-of-the-art approaches in language-driven object navigation. We further demonstrate its practicality through real-world deployment on a quadruped robot, achieving robust and effective navigation performance.
format Preprint
id arxiv_https___arxiv_org_abs_2508_05021
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MAG-Nav: Language-Driven Object Navigation Leveraging Memory-Reserved Active Grounding
Zhang, Weifan
Li, Tingguang
Liu, Yuzhen
Robotics
Visual navigation in unknown environments based solely on natural language descriptions is a key capability for intelligent robots. In this work, we propose a navigation framework built upon off-the-shelf Visual Language Models (VLMs), enhanced with two human-inspired mechanisms: perspective-based active grounding, which dynamically adjusts the robot's viewpoint for improved visual inspection, and historical memory backtracking, which enables the system to retain and re-evaluate uncertain observations over time. Unlike existing approaches that passively rely on incidental visual inputs, our method actively optimizes perception and leverages memory to resolve ambiguity, significantly improving vision-language grounding in complex, unseen environments. Our framework operates in a zero-shot manner, achieving strong generalization to diverse and open-ended language descriptions without requiring labeled data or model fine-tuning. Experimental results on Habitat-Matterport 3D (HM3D) show that our method outperforms state-of-the-art approaches in language-driven object navigation. We further demonstrate its practicality through real-world deployment on a quadruped robot, achieving robust and effective navigation performance.
title MAG-Nav: Language-Driven Object Navigation Leveraging Memory-Reserved Active Grounding
topic Robotics
url https://arxiv.org/abs/2508.05021