AwareVLN: Reasoning with Self-awareness for Vision-Language Navigation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Guo, Wenxuan, Xu, Xiuwei, Liu, Yichen, Li, Xiangyu, Yin, Hang, Chen, Huangxing, Zheng, Wenzhao, Feng, Jianjiang, Zhou, Jie, Lu, Jiwen
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913153428750336
author Guo, Wenxuan
Xu, Xiuwei
Liu, Yichen
Li, Xiangyu
Yin, Hang
Chen, Huangxing
Zheng, Wenzhao
Feng, Jianjiang
Zhou, Jie
Lu, Jiwen
author_facet Guo, Wenxuan
Xu, Xiuwei
Liu, Yichen
Li, Xiangyu
Yin, Hang
Chen, Huangxing
Zheng, Wenzhao
Feng, Jianjiang
Zhou, Jie
Lu, Jiwen
contents Vision-and-Language Navigation (VLN) requires an agent to ground language instructions to its own movement within a visual environment. While state-of-the-art methods leverage the reasoning capabilities of Vision-Language Models (VLMs) for end-to-end action prediction, they often lack an explicit and explainable understanding of the relationships between the agent, the instruction, and the scene. Conversely, explicitly building a scene map for heuristic planning is intuitively appealing but relies on additional 3D sensors and hinders large-scale vision-language pre-training. To bridge this gap, we propose AwareVLN, a novel framework that equips the navigation model with a self-aware reasoning mechanism, enabling it to understand the agent's state and task progress in a fully end-to-end and data-driven manner. Our approach features two key innovations: (1) a structural reasoning module that fosters spatial and task-oriented self-awareness, and (2) an automatic data engine with progress division for effective training. Extensive experiments on various datasets in Habitat simulator show our AwareVLN significantly outperforms previous state-of-the-art vision-language navigation methods. Project page: https://gwxuan.github.io/AwareVLN/.
format Preprint
id arxiv_https___arxiv_org_abs_2605_22816
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle AwareVLN: Reasoning with Self-awareness for Vision-Language Navigation
Guo, Wenxuan
Xu, Xiuwei
Liu, Yichen
Li, Xiangyu
Yin, Hang
Chen, Huangxing
Zheng, Wenzhao
Feng, Jianjiang
Zhou, Jie
Lu, Jiwen
Robotics
Computer Vision and Pattern Recognition
Vision-and-Language Navigation (VLN) requires an agent to ground language instructions to its own movement within a visual environment. While state-of-the-art methods leverage the reasoning capabilities of Vision-Language Models (VLMs) for end-to-end action prediction, they often lack an explicit and explainable understanding of the relationships between the agent, the instruction, and the scene. Conversely, explicitly building a scene map for heuristic planning is intuitively appealing but relies on additional 3D sensors and hinders large-scale vision-language pre-training. To bridge this gap, we propose AwareVLN, a novel framework that equips the navigation model with a self-aware reasoning mechanism, enabling it to understand the agent's state and task progress in a fully end-to-end and data-driven manner. Our approach features two key innovations: (1) a structural reasoning module that fosters spatial and task-oriented self-awareness, and (2) an automatic data engine with progress division for effective training. Extensive experiments on various datasets in Habitat simulator show our AwareVLN significantly outperforms previous state-of-the-art vision-language navigation methods. Project page: https://gwxuan.github.io/AwareVLN/.
title AwareVLN: Reasoning with Self-awareness for Vision-Language Navigation
topic Robotics
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.22816