GeoNav: Empowering MLLMs with dual-scale geospatial reasoning for language-goal aerial navigation

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Xu, Haotian, Hu, Yue, Gao, Chen, Zhu, Zhengqiu, Zhao, Yong, Li, Yong, Yin, Quanjun
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914376891498496
author Xu, Haotian
Hu, Yue
Gao, Chen
Zhu, Zhengqiu
Zhao, Yong
Li, Yong
Yin, Quanjun
author_facet Xu, Haotian
Hu, Yue
Gao, Chen
Zhu, Zhengqiu
Zhao, Yong
Li, Yong
Yin, Quanjun
contents Language-goal aerial navigation requires UAVs to localize targets in the complex outdoors, such as urban blocks based on textual instructions. The indoor methods are often hard to scale to urban scenes due to ambiguous objects, limited visual field, and spatial reasoning. In this work, we propose GeoNav, a multi-modal agent for long-range aerial navigation with geospatial awareness. GeoNav operates in three phases-landmark navigation, target search, and precise localization-mimicking human coarse-to-fine spatial reasoning patterns. To support such reasoning, it dynamically builds dual-scale spatial representations. The first is a global but schematic cognitive map, which fuses prior geographic knowledge and embodied visual cues into a top-down and explicit annotated form. It enables fast navigation to the landmark region via intuitive map-based reasoning. The second is a local but delicate scene graph representing hierarchical spatial relationships between landmarks and objects, utilized for accurate target localization. On top of the structured memory, GeoNav employs a spatial chain-of-thought mechanism to enable MLLMs with efficient and interpretable action-making across stages. On the CityNav benchmark, GeoNav surpasses the current SOTA up to 18.4% in success rate and significantly eliminates navigation error. The ablation studies highlight the importance of each module, positioning structured spatial perception as the key to advanced UAV navigation. Published in Pattern Recognition, 2026.
format Preprint
id arxiv_https___arxiv_org_abs_2504_09587
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle GeoNav: Empowering MLLMs with dual-scale geospatial reasoning for language-goal aerial navigation
Xu, Haotian
Hu, Yue
Gao, Chen
Zhu, Zhengqiu
Zhao, Yong
Li, Yong
Yin, Quanjun
Robotics
Language-goal aerial navigation requires UAVs to localize targets in the complex outdoors, such as urban blocks based on textual instructions. The indoor methods are often hard to scale to urban scenes due to ambiguous objects, limited visual field, and spatial reasoning. In this work, we propose GeoNav, a multi-modal agent for long-range aerial navigation with geospatial awareness. GeoNav operates in three phases-landmark navigation, target search, and precise localization-mimicking human coarse-to-fine spatial reasoning patterns. To support such reasoning, it dynamically builds dual-scale spatial representations. The first is a global but schematic cognitive map, which fuses prior geographic knowledge and embodied visual cues into a top-down and explicit annotated form. It enables fast navigation to the landmark region via intuitive map-based reasoning. The second is a local but delicate scene graph representing hierarchical spatial relationships between landmarks and objects, utilized for accurate target localization. On top of the structured memory, GeoNav employs a spatial chain-of-thought mechanism to enable MLLMs with efficient and interpretable action-making across stages. On the CityNav benchmark, GeoNav surpasses the current SOTA up to 18.4% in success rate and significantly eliminates navigation error. The ablation studies highlight the importance of each module, positioning structured spatial perception as the key to advanced UAV navigation. Published in Pattern Recognition, 2026.
title GeoNav: Empowering MLLMs with dual-scale geospatial reasoning for language-goal aerial navigation
topic Robotics
url https://arxiv.org/abs/2504.09587