NavSpace: How Navigation Agents Follow Spatial Intelligence Instructions

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Yang, Haolin, Long, Yuxing, Yu, Zhuoyuan, Yang, Zihan, Wang, Minghan, Xu, Jiapeng, Wang, Yihan, Yu, Ziyan, Cai, Wenzhe, Kang, Lei, Dong, Hao
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866910046841995264
author Yang, Haolin
Long, Yuxing
Yu, Zhuoyuan
Yang, Zihan
Wang, Minghan
Xu, Jiapeng
Wang, Yihan
Yu, Ziyan
Cai, Wenzhe
Kang, Lei
Dong, Hao
author_facet Yang, Haolin
Long, Yuxing
Yu, Zhuoyuan
Yang, Zihan
Wang, Minghan
Xu, Jiapeng
Wang, Yihan
Yu, Ziyan
Cai, Wenzhe
Kang, Lei
Dong, Hao
contents Instruction-following navigation is a key step toward embodied intelligence. Prior benchmarks mainly focus on semantic understanding but overlook systematically evaluating navigation agents' spatial perception and reasoning capabilities. In this work, we introduce the NavSpace benchmark, which contains six task categories and 1,228 trajectory-instruction pairs designed to probe the spatial intelligence of navigation agents. On this benchmark, we comprehensively evaluate 22 navigation agents, including state-of-the-art navigation models and multimodal large language models. The evaluation results lift the veil on spatial intelligence in embodied navigation. Furthermore, we propose SNav, a new spatially intelligent navigation model. SNav outperforms existing navigation agents on NavSpace and real robot tests, establishing a strong baseline for future work.
format Preprint
id arxiv_https___arxiv_org_abs_2510_08173
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle NavSpace: How Navigation Agents Follow Spatial Intelligence Instructions
Yang, Haolin
Long, Yuxing
Yu, Zhuoyuan
Yang, Zihan
Wang, Minghan
Xu, Jiapeng
Wang, Yihan
Yu, Ziyan
Cai, Wenzhe
Kang, Lei
Dong, Hao
Robotics
Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
Instruction-following navigation is a key step toward embodied intelligence. Prior benchmarks mainly focus on semantic understanding but overlook systematically evaluating navigation agents' spatial perception and reasoning capabilities. In this work, we introduce the NavSpace benchmark, which contains six task categories and 1,228 trajectory-instruction pairs designed to probe the spatial intelligence of navigation agents. On this benchmark, we comprehensively evaluate 22 navigation agents, including state-of-the-art navigation models and multimodal large language models. The evaluation results lift the veil on spatial intelligence in embodied navigation. Furthermore, we propose SNav, a new spatially intelligent navigation model. SNav outperforms existing navigation agents on NavSpace and real robot tests, establishing a strong baseline for future work.
title NavSpace: How Navigation Agents Follow Spatial Intelligence Instructions
topic Robotics
Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.08173