Uncertainty-Aware Gaussian Map for Vision-Language Navigation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gao, Jianzhe, Liu, Rui, Xu, Yuxuan, Cao, Tongtong, Zhang, Yingxue, Zhang, Zhanguang, Peng, Sida, Yang, Yi, Wang, Wenguan
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910257930829824
author Gao, Jianzhe
Liu, Rui
Xu, Yuxuan
Cao, Tongtong
Zhang, Yingxue
Zhang, Zhanguang
Peng, Sida
Yang, Yi
Wang, Wenguan
author_facet Gao, Jianzhe
Liu, Rui
Xu, Yuxuan
Cao, Tongtong
Zhang, Yingxue
Zhang, Zhanguang
Peng, Sida
Yang, Yi
Wang, Wenguan
contents Vision-Language Navigation (VLN) requires an agent to navigate 3D environments following natural language instructions. During navigation, existing agents commonly encounter perceptual uncertainty, such as insufficient evidence for reliable grounding or ambiguity in interpreting spatial cues, yet they typically ignore such information when predicting actions. In this work, we explicitly model three forms of perceptual uncertainty (i.e., geometric, semantic, and appearance uncertainty) and integrate them into the agent's observation space to enable informed decision-making. Concretely, our agent first constructs a Semantic Gaussian Map (SGM), composed of differentiable 3D Gaussian primitives initialized from panoramic observations, that encodes both the geometric structure and semantic content of the environment. On top of SGM, geometric uncertainty is estimated through variational perturbations of Gaussian position and scale to assess structural reliability; semantic uncertainty is captured by perturbing Gaussian semantic attributes to reveal ambiguous interpretations; and appearance uncertainty is characterized by Fisher Information, which measures the sensitivity of rendered observations to Gaussian-level variations. These uncertainties are incorporated into SGM, extending it into a unified 3D Value Map, which grounds them as affordances and constraints that support reliable navigation. Comprehensive evaluations across multiple VLN benchmarks show the effectiveness of our agent.
format Preprint
id arxiv_https___arxiv_org_abs_2605_26503
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Uncertainty-Aware Gaussian Map for Vision-Language Navigation
Gao, Jianzhe
Liu, Rui
Xu, Yuxuan
Cao, Tongtong
Zhang, Yingxue
Zhang, Zhanguang
Peng, Sida
Yang, Yi
Wang, Wenguan
Computer Vision and Pattern Recognition
Vision-Language Navigation (VLN) requires an agent to navigate 3D environments following natural language instructions. During navigation, existing agents commonly encounter perceptual uncertainty, such as insufficient evidence for reliable grounding or ambiguity in interpreting spatial cues, yet they typically ignore such information when predicting actions. In this work, we explicitly model three forms of perceptual uncertainty (i.e., geometric, semantic, and appearance uncertainty) and integrate them into the agent's observation space to enable informed decision-making. Concretely, our agent first constructs a Semantic Gaussian Map (SGM), composed of differentiable 3D Gaussian primitives initialized from panoramic observations, that encodes both the geometric structure and semantic content of the environment. On top of SGM, geometric uncertainty is estimated through variational perturbations of Gaussian position and scale to assess structural reliability; semantic uncertainty is captured by perturbing Gaussian semantic attributes to reveal ambiguous interpretations; and appearance uncertainty is characterized by Fisher Information, which measures the sensitivity of rendered observations to Gaussian-level variations. These uncertainties are incorporated into SGM, extending it into a unified 3D Value Map, which grounds them as affordances and constraints that support reliable navigation. Comprehensive evaluations across multiple VLN benchmarks show the effectiveness of our agent.
title Uncertainty-Aware Gaussian Map for Vision-Language Navigation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.26503