AutoFly: Vision-Language-Action Model for UAV Autonomous Navigation in the Wild

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Sun, Xiaolou, Si, Wufei, Ni, Wenhui, Li, Yuntian, Wu, Dongming, Xie, Fei, Guan, Runwei, Xu, He-Yang, Ding, Henghui, Wu, Yuan, Yue, Yutao, Huang, Yongming, Xiong, Hui
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866918331078934528
author Sun, Xiaolou
Si, Wufei
Ni, Wenhui
Li, Yuntian
Wu, Dongming
Xie, Fei
Guan, Runwei
Xu, He-Yang
Ding, Henghui
Wu, Yuan
Yue, Yutao
Huang, Yongming
Xiong, Hui
author_facet Sun, Xiaolou
Si, Wufei
Ni, Wenhui
Li, Yuntian
Wu, Dongming
Xie, Fei
Guan, Runwei
Xu, He-Yang
Ding, Henghui
Wu, Yuan
Yue, Yutao
Huang, Yongming
Xiong, Hui
contents Vision-language navigation (VLN) requires intelligent agents to navigate environments by interpreting linguistic instructions alongside visual observations, serving as a cornerstone task in Embodied AI. Current VLN research for unmanned aerial vehicles (UAVs) relies on detailed, pre-specified instructions to guide the UAV along predetermined routes. However, real-world outdoor exploration typically occurs in unknown environments where detailed navigation instructions are unavailable. Instead, only coarse-grained positional or directional guidance can be provided, requiring UAVs to autonomously navigate through continuous planning and obstacle avoidance. To bridge this gap, we propose AutoFly, an end-to-end Vision-Language-Action (VLA) model for autonomous UAV navigation. AutoFly incorporates a pseudo-depth encoder that derives depth-aware features from RGB inputs to enhance spatial reasoning, coupled with a progressive two-stage training strategy that effectively aligns visual, depth, and linguistic representations with action policies. Moreover, existing VLN datasets have fundamental limitations for real-world autonomous navigation, stemming from their heavy reliance on explicit instruction-following over autonomous decision-making and insufficient real-world data. To address these issues, we construct a novel autonomous navigation dataset that shifts the paradigm from instruction-following to autonomous behavior modeling through: (1) trajectory collection emphasizing continuous obstacle avoidance, autonomous planning, and recognition workflows; (2) comprehensive real-world data integration. Experimental results demonstrate that AutoFly achieves a 3.9% higher success rate compared to state-of-the-art VLA baselines, with consistent performance across simulated and real environments.
format Preprint
id arxiv_https___arxiv_org_abs_2602_09657
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle AutoFly: Vision-Language-Action Model for UAV Autonomous Navigation in the Wild
Sun, Xiaolou
Si, Wufei
Ni, Wenhui
Li, Yuntian
Wu, Dongming
Xie, Fei
Guan, Runwei
Xu, He-Yang
Ding, Henghui
Wu, Yuan
Yue, Yutao
Huang, Yongming
Xiong, Hui
Robotics
Vision-language navigation (VLN) requires intelligent agents to navigate environments by interpreting linguistic instructions alongside visual observations, serving as a cornerstone task in Embodied AI. Current VLN research for unmanned aerial vehicles (UAVs) relies on detailed, pre-specified instructions to guide the UAV along predetermined routes. However, real-world outdoor exploration typically occurs in unknown environments where detailed navigation instructions are unavailable. Instead, only coarse-grained positional or directional guidance can be provided, requiring UAVs to autonomously navigate through continuous planning and obstacle avoidance. To bridge this gap, we propose AutoFly, an end-to-end Vision-Language-Action (VLA) model for autonomous UAV navigation. AutoFly incorporates a pseudo-depth encoder that derives depth-aware features from RGB inputs to enhance spatial reasoning, coupled with a progressive two-stage training strategy that effectively aligns visual, depth, and linguistic representations with action policies. Moreover, existing VLN datasets have fundamental limitations for real-world autonomous navigation, stemming from their heavy reliance on explicit instruction-following over autonomous decision-making and insufficient real-world data. To address these issues, we construct a novel autonomous navigation dataset that shifts the paradigm from instruction-following to autonomous behavior modeling through: (1) trajectory collection emphasizing continuous obstacle avoidance, autonomous planning, and recognition workflows; (2) comprehensive real-world data integration. Experimental results demonstrate that AutoFly achieves a 3.9% higher success rate compared to state-of-the-art VLA baselines, with consistent performance across simulated and real environments.
title AutoFly: Vision-Language-Action Model for UAV Autonomous Navigation in the Wild
topic Robotics
url https://arxiv.org/abs/2602.09657