SpatialNav: Leveraging Spatial Scene Graphs for Zero-Shot Vision-and-Language Navigation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Jiwen, Li, Zejun, Wang, Siyuan, Shi, Xiangyu, Wei, Zhongyu, Wu, Qi
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912814959951872
author Zhang, Jiwen
Li, Zejun
Wang, Siyuan
Shi, Xiangyu
Wei, Zhongyu
Wu, Qi
author_facet Zhang, Jiwen
Li, Zejun
Wang, Siyuan
Shi, Xiangyu
Wei, Zhongyu
Wu, Qi
contents Although learning-based vision-and-language navigation (VLN) agents can learn spatial knowledge implicitly from large-scale training data, zero-shot VLN agents lack this process, relying primarily on local observations for navigation, which leads to inefficient exploration and a significant performance gap. To deal with the problem, we consider a zero-shot VLN setting that agents are allowed to fully explore the environment before task execution. Then, we construct the Spatial Scene Graph (SSG) to explicitly capture global spatial structure and semantics in the explored environment. Based on the SSG, we introduce SpatialNav, a zero-shot VLN agent that integrates an agent-centric spatial map, a compass-aligned visual representation, and a remote object localization strategy for efficient navigation. Comprehensive experiments in both discrete and continuous environments demonstrate that SpatialNav significantly outperforms existing zero-shot agents and clearly narrows the gap with state-of-the-art learning-based methods. Such results highlight the importance of global spatial representations for generalizable navigation.
format Preprint
id arxiv_https___arxiv_org_abs_2601_06806
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle SpatialNav: Leveraging Spatial Scene Graphs for Zero-Shot Vision-and-Language Navigation
Zhang, Jiwen
Li, Zejun
Wang, Siyuan
Shi, Xiangyu
Wei, Zhongyu
Wu, Qi
Computer Vision and Pattern Recognition
Artificial Intelligence
Robotics
Although learning-based vision-and-language navigation (VLN) agents can learn spatial knowledge implicitly from large-scale training data, zero-shot VLN agents lack this process, relying primarily on local observations for navigation, which leads to inefficient exploration and a significant performance gap. To deal with the problem, we consider a zero-shot VLN setting that agents are allowed to fully explore the environment before task execution. Then, we construct the Spatial Scene Graph (SSG) to explicitly capture global spatial structure and semantics in the explored environment. Based on the SSG, we introduce SpatialNav, a zero-shot VLN agent that integrates an agent-centric spatial map, a compass-aligned visual representation, and a remote object localization strategy for efficient navigation. Comprehensive experiments in both discrete and continuous environments demonstrate that SpatialNav significantly outperforms existing zero-shot agents and clearly narrows the gap with state-of-the-art learning-based methods. Such results highlight the importance of global spatial representations for generalizable navigation.
title SpatialNav: Leveraging Spatial Scene Graphs for Zero-Shot Vision-and-Language Navigation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Robotics
url https://arxiv.org/abs/2601.06806