AdaNav: Adaptive Reasoning with Uncertainty for Vision-Language Navigation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ding, Xin, Wei, Jianyu, Yang, Yifan, Jiang, Shiqi, Zhang, Qianxi, Wu, Hao, Jia, Fucheng, Mi, Liang, Yan, Yuxuan, Wang, Weijun, Liu, Yunxin, Chen, Zhibo, Cao, Ting
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911183538225152
author Ding, Xin
Wei, Jianyu
Yang, Yifan
Jiang, Shiqi
Zhang, Qianxi
Wu, Hao
Jia, Fucheng
Mi, Liang
Yan, Yuxuan
Wang, Weijun
Liu, Yunxin
Chen, Zhibo
Cao, Ting
author_facet Ding, Xin
Wei, Jianyu
Yang, Yifan
Jiang, Shiqi
Zhang, Qianxi
Wu, Hao
Jia, Fucheng
Mi, Liang
Yan, Yuxuan
Wang, Weijun
Liu, Yunxin
Chen, Zhibo
Cao, Ting
contents Vision Language Navigation (VLN) requires agents to follow natural language instructions by grounding them in sequential visual observations over long horizons. Explicit reasoning could enhance temporal consistency and perception action alignment, but reasoning at fixed steps often leads to suboptimal performance and unnecessary computation. To address this, we propose AdaNav, an uncertainty-based adaptive reasoning framework for VLN. At its core is the Uncertainty Adaptive Reasoning Block (UAR), a lightweight plugin that dynamically triggers reasoning. We introduce Action Entropy as a policy prior for UAR and progressively refine it through a Heuristics to RL training method, enabling agents to learn difficulty aware reasoning policies under the strict data limitations of embodied tasks. Results show that with only 6K training samples, AdaNav achieves substantial gains over closed source models trained on million scale data, improving success rate by 20% on R2R val-unseen, 11.7% on RxR-CE, and 11.4% in real world scenes. The code is available at https://github.com/xinding-sys/AdaNav.
format Preprint
id arxiv_https___arxiv_org_abs_2509_24387
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle AdaNav: Adaptive Reasoning with Uncertainty for Vision-Language Navigation
Ding, Xin
Wei, Jianyu
Yang, Yifan
Jiang, Shiqi
Zhang, Qianxi
Wu, Hao
Jia, Fucheng
Mi, Liang
Yan, Yuxuan
Wang, Weijun
Liu, Yunxin
Chen, Zhibo
Cao, Ting
Robotics
Vision Language Navigation (VLN) requires agents to follow natural language instructions by grounding them in sequential visual observations over long horizons. Explicit reasoning could enhance temporal consistency and perception action alignment, but reasoning at fixed steps often leads to suboptimal performance and unnecessary computation. To address this, we propose AdaNav, an uncertainty-based adaptive reasoning framework for VLN. At its core is the Uncertainty Adaptive Reasoning Block (UAR), a lightweight plugin that dynamically triggers reasoning. We introduce Action Entropy as a policy prior for UAR and progressively refine it through a Heuristics to RL training method, enabling agents to learn difficulty aware reasoning policies under the strict data limitations of embodied tasks. Results show that with only 6K training samples, AdaNav achieves substantial gains over closed source models trained on million scale data, improving success rate by 20% on R2R val-unseen, 11.7% on RxR-CE, and 11.4% in real world scenes. The code is available at https://github.com/xinding-sys/AdaNav.
title AdaNav: Adaptive Reasoning with Uncertainty for Vision-Language Navigation
topic Robotics
url https://arxiv.org/abs/2509.24387