Grounded Vision-Language Navigation for UAVs with Open-Vocabulary Goal Understanding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Yuhang, Yu, Haosheng, Xiao, Jiaping, Feroskhan, Mir
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915339613241344
author Zhang, Yuhang
Yu, Haosheng
Xiao, Jiaping
Feroskhan, Mir
author_facet Zhang, Yuhang
Yu, Haosheng
Xiao, Jiaping
Feroskhan, Mir
contents Vision-and-language navigation (VLN) is a long-standing challenge in autonomous robotics, aiming to empower agents with the ability to follow human instructions while navigating complex environments. Two key bottlenecks remain in this field: generalization to out-of-distribution environments and reliance on fixed discrete action spaces. To address these challenges, we propose Vision-Language Fly (VLFly), a framework tailored for Unmanned Aerial Vehicles (UAVs) to execute language-guided flight. Without the requirement for localization or active ranging sensors, VLFly outputs continuous velocity commands purely from egocentric observations captured by an onboard monocular camera. The VLFly integrates three modules: an instruction encoder based on a large language model (LLM) that reformulates high-level language into structured prompts, a goal retriever powered by a vision-language model (VLM) that matches these prompts to goal images via vision-language similarity, and a waypoint planner that generates executable trajectories for real-time UAV control. VLFly is evaluated across diverse simulation environments without additional fine-tuning and consistently outperforms all baselines. Moreover, real-world VLN tasks in indoor and outdoor environments under direct and indirect instructions demonstrate that VLFly achieves robust open-vocabulary goal understanding and generalized navigation capabilities, even in the presence of abstract language input.
format Preprint
id arxiv_https___arxiv_org_abs_2506_10756
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Grounded Vision-Language Navigation for UAVs with Open-Vocabulary Goal Understanding
Zhang, Yuhang
Yu, Haosheng
Xiao, Jiaping
Feroskhan, Mir
Robotics
Artificial Intelligence
Vision-and-language navigation (VLN) is a long-standing challenge in autonomous robotics, aiming to empower agents with the ability to follow human instructions while navigating complex environments. Two key bottlenecks remain in this field: generalization to out-of-distribution environments and reliance on fixed discrete action spaces. To address these challenges, we propose Vision-Language Fly (VLFly), a framework tailored for Unmanned Aerial Vehicles (UAVs) to execute language-guided flight. Without the requirement for localization or active ranging sensors, VLFly outputs continuous velocity commands purely from egocentric observations captured by an onboard monocular camera. The VLFly integrates three modules: an instruction encoder based on a large language model (LLM) that reformulates high-level language into structured prompts, a goal retriever powered by a vision-language model (VLM) that matches these prompts to goal images via vision-language similarity, and a waypoint planner that generates executable trajectories for real-time UAV control. VLFly is evaluated across diverse simulation environments without additional fine-tuning and consistently outperforms all baselines. Moreover, real-world VLN tasks in indoor and outdoor environments under direct and indirect instructions demonstrate that VLFly achieves robust open-vocabulary goal understanding and generalized navigation capabilities, even in the presence of abstract language input.
title Grounded Vision-Language Navigation for UAVs with Open-Vocabulary Goal Understanding
topic Robotics
Artificial Intelligence
url https://arxiv.org/abs/2506.10756