History-Enhanced Two-Stage Transformer for Aerial Vision-and-Language Navigation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ding, Xichen, Gao, Jianzhe, Pan, Cong, Wang, Wenguan, Qin, Jie
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914205106438144
author Ding, Xichen
Gao, Jianzhe
Pan, Cong
Wang, Wenguan
Qin, Jie
author_facet Ding, Xichen
Gao, Jianzhe
Pan, Cong
Wang, Wenguan
Qin, Jie
contents Aerial Vision-and-Language Navigation (AVLN) requires Unmanned Aerial Vehicle (UAV) agents to localize targets in large-scale urban environments based on linguistic instructions. While successful navigation demands both global environmental reasoning and local scene comprehension, existing UAV agents typically adopt mono-granularity frameworks that struggle to balance these two aspects. To address this limitation, this work proposes a History-Enhanced Two-Stage Transformer (HETT) framework, which integrates the two aspects through a coarse-to-fine navigation pipeline. Specifically, HETT first predicts coarse-grained target positions by fusing spatial landmarks and historical context, then refines actions via fine-grained visual analysis. In addition, a historical grid map is designed to dynamically aggregate visual features into a structured spatial memory, enhancing comprehensive scene awareness. Additionally, the CityNav dataset annotations are manually refined to enhance data quality. Experiments on the refined CityNav dataset show that HETT delivers significant performance gains, while extensive ablation studies further verify the effectiveness of each component.
format Preprint
id arxiv_https___arxiv_org_abs_2512_14222
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle History-Enhanced Two-Stage Transformer for Aerial Vision-and-Language Navigation
Ding, Xichen
Gao, Jianzhe
Pan, Cong
Wang, Wenguan
Qin, Jie
Computer Vision and Pattern Recognition
Robotics
Aerial Vision-and-Language Navigation (AVLN) requires Unmanned Aerial Vehicle (UAV) agents to localize targets in large-scale urban environments based on linguistic instructions. While successful navigation demands both global environmental reasoning and local scene comprehension, existing UAV agents typically adopt mono-granularity frameworks that struggle to balance these two aspects. To address this limitation, this work proposes a History-Enhanced Two-Stage Transformer (HETT) framework, which integrates the two aspects through a coarse-to-fine navigation pipeline. Specifically, HETT first predicts coarse-grained target positions by fusing spatial landmarks and historical context, then refines actions via fine-grained visual analysis. In addition, a historical grid map is designed to dynamically aggregate visual features into a structured spatial memory, enhancing comprehensive scene awareness. Additionally, the CityNav dataset annotations are manually refined to enhance data quality. Experiments on the refined CityNav dataset show that HETT delivers significant performance gains, while extensive ablation studies further verify the effectiveness of each component.
title History-Enhanced Two-Stage Transformer for Aerial Vision-and-Language Navigation
topic Computer Vision and Pattern Recognition
Robotics
url https://arxiv.org/abs/2512.14222