ProFocus: Proactive Perception and Focused Reasoning in Vision-and-Language Navigation

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Xue, Wei, Li, Mingcheng, Wu, Xuecheng, Tang, Jingqun, Yang, Dingkang, Zhang, Lihua
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866917340190343168
author Xue, Wei
Li, Mingcheng
Wu, Xuecheng
Tang, Jingqun
Yang, Dingkang
Zhang, Lihua
author_facet Xue, Wei
Li, Mingcheng
Wu, Xuecheng
Tang, Jingqun
Yang, Dingkang
Zhang, Lihua
contents Vision-and-Language Navigation (VLN) requires agents to accurately perceive complex visual environments and reason over navigation instructions and histories. However, existing methods passively process redundant visual inputs and treat all historical contexts indiscriminately, resulting in inefficient perception and unfocused reasoning. To address these challenges, we propose \textbf{ProFocus}, a training-free progressive framework that unifies \underline{Pro}active Perception and \underline{Focus}ed Reasoning through collaboration between large language models (LLMs) and vision-language models (VLMs). For proactive perception, ProFocus transforms panoramic observations into structured ego-centric semantic maps, enabling the orchestration agent to identify missing visual information needed for reliable decision-making, and to generate targeted visual queries with corresponding focus regions that guide the perception agent to acquire the required observations. For focused reasoning, we propose Branch-Diverse Monte Carlo Tree Search (BD-MCTS) to identify top-$k$ high-value waypoints from extensive historical candidates. The decision agent focuses reasoning on the historical contexts associated with these waypoints, rather than considering all historical waypoints equally. Extensive experiments validate the effectiveness of ProFocus, achieving state-of-the-art performance among zero-shot methods on R2R and REVERIE benchmarks.
format Preprint
id arxiv_https___arxiv_org_abs_2603_05530
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle ProFocus: Proactive Perception and Focused Reasoning in Vision-and-Language Navigation
Xue, Wei
Li, Mingcheng
Wu, Xuecheng
Tang, Jingqun
Yang, Dingkang
Zhang, Lihua
Robotics
Computer Vision and Pattern Recognition
Vision-and-Language Navigation (VLN) requires agents to accurately perceive complex visual environments and reason over navigation instructions and histories. However, existing methods passively process redundant visual inputs and treat all historical contexts indiscriminately, resulting in inefficient perception and unfocused reasoning. To address these challenges, we propose \textbf{ProFocus}, a training-free progressive framework that unifies \underline{Pro}active Perception and \underline{Focus}ed Reasoning through collaboration between large language models (LLMs) and vision-language models (VLMs). For proactive perception, ProFocus transforms panoramic observations into structured ego-centric semantic maps, enabling the orchestration agent to identify missing visual information needed for reliable decision-making, and to generate targeted visual queries with corresponding focus regions that guide the perception agent to acquire the required observations. For focused reasoning, we propose Branch-Diverse Monte Carlo Tree Search (BD-MCTS) to identify top-$k$ high-value waypoints from extensive historical candidates. The decision agent focuses reasoning on the historical contexts associated with these waypoints, rather than considering all historical waypoints equally. Extensive experiments validate the effectiveness of ProFocus, achieving state-of-the-art performance among zero-shot methods on R2R and REVERIE benchmarks.
title ProFocus: Proactive Perception and Focused Reasoning in Vision-and-Language Navigation
topic Robotics
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2603.05530