VLN-Pilot: Large Vision-Language Model as an Autonomous Indoor Drone Operator
Fuente:
arXiv
Saved in:
| Main Authors: | Dominguez-Dager, Bessie, Suescun-Ferrandiz, Sergio, Escalona, Felix, Gomez-Donoso, Francisco, Cazorla, Miguel |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CHIRLA: Comprehensive High-resolution Identification and Re-identification for Large-scale Analysis
by: Dominguez-Dager, Bessie, et al.
Published: (2025)
by: Dominguez-Dager, Bessie, et al.
Published: (2025)
AIDEN: Design and Pilot Study of an AI Assistant for the Visually Impaired
by: Marquez-Carpintero, Luis, et al.
Published: (2025)
by: Marquez-Carpintero, Luis, et al.
Published: (2025)
CADDI: An in-Class Activity Detection Dataset using IMU data from low-cost sensors
by: Marquez-Carpintero, Luis, et al.
Published: (2025)
by: Marquez-Carpintero, Luis, et al.
Published: (2025)
AdaVLN: Towards Visual Language Navigation in Continuous Indoor Environments with Moving Humans
by: Loh, Dillon, et al.
Published: (2024)
by: Loh, Dillon, et al.
Published: (2024)
AwareVLN: Reasoning with Self-awareness for Vision-Language Navigation
by: Guo, Wenxuan, et al.
Published: (2026)
by: Guo, Wenxuan, et al.
Published: (2026)
FlexVLN: Flexible Adaptation for Diverse Vision-and-Language Navigation Tasks
by: Zhang, Siqi, et al.
Published: (2025)
by: Zhang, Siqi, et al.
Published: (2025)
UAV-VLN: End-to-End Vision Language guided Navigation for UAVs
by: Saxena, Pranav, et al.
Published: (2025)
by: Saxena, Pranav, et al.
Published: (2025)
AgriVLN: Vision-and-Language Navigation for Agricultural Robots
by: Zhao, Xiaobei, et al.
Published: (2025)
by: Zhao, Xiaobei, et al.
Published: (2025)
FantasyVLN: Unified Multimodal Chain-of-Thought Reasoning for Vision-Language Navigation
by: Zuo, Jing, et al.
Published: (2026)
by: Zuo, Jing, et al.
Published: (2026)
WorldVLN: Autoregressive World Action Model for Aerial Vision-Language Navigation
by: Zhao, Baining, et al.
Published: (2026)
by: Zhao, Baining, et al.
Published: (2026)
GC-VLN: Instruction as Graph Constraints for Training-free Vision-and-Language Navigation
by: Yin, Hang, et al.
Published: (2025)
by: Yin, Hang, et al.
Published: (2025)
StreamVLN: Streaming Vision-and-Language Navigation via SlowFast Context Modeling
by: Wei, Meng, et al.
Published: (2025)
by: Wei, Meng, et al.
Published: (2025)
JanusVLN: Decoupling Semantics and Spatiality with Dual Implicit Memory for Vision-Language Navigation
by: Zeng, Shuang, et al.
Published: (2025)
by: Zeng, Shuang, et al.
Published: (2025)
SE-VLN: A Self-Evolving Vision-Language Navigation Framework Based on Multimodal Large Language Models
by: Dong, Xiangyu, et al.
Published: (2025)
by: Dong, Xiangyu, et al.
Published: (2025)
VLN-NF: Feasibility-Aware Vision-and-Language Navigation with False-Premise Instructions
by: Su, Hung-Ting, et al.
Published: (2026)
by: Su, Hung-Ting, et al.
Published: (2026)
Semantic-Aware Guided Drone Exploration for Language-Conditioned 3D Indoor Mapping
by: Vegesna, Nitin, et al.
Published: (2026)
by: Vegesna, Nitin, et al.
Published: (2026)
HiMemVLN: Enhancing Reliability of Open-Source Zero-Shot Vision-and-Language Navigation with Hierarchical Memory System
by: Lyu, Kailin, et al.
Published: (2026)
by: Lyu, Kailin, et al.
Published: (2026)
VLN-MME: Diagnosing MLLMs as Language-guided Visual Navigation agents
by: Zhao, Xunyi, et al.
Published: (2025)
by: Zhao, Xunyi, et al.
Published: (2025)
ActiveVLN: Towards Active Exploration via Multi-Turn RL in Vision-and-Language Navigation
by: Zhang, Zekai, et al.
Published: (2025)
by: Zhang, Zekai, et al.
Published: (2025)
Learning on the Fly: Replay-Based Continual Object Perception for Indoor Drones
by: Nae, Sebastian-Ion, et al.
Published: (2026)
by: Nae, Sebastian-Ion, et al.
Published: (2026)
VISTAv2: World Imagination for Indoor Vision-and-Language Navigation
by: Huang, Yanjia, et al.
Published: (2025)
by: Huang, Yanjia, et al.
Published: (2025)
CLIPSwarm: Generating Drone Shows from Text Prompts with Vision-Language Models
by: Pueyo, Pablo, et al.
Published: (2024)
by: Pueyo, Pablo, et al.
Published: (2024)
Senna: Bridging Large Vision-Language Models and End-to-End Autonomous Driving
by: Jiang, Bo, et al.
Published: (2024)
by: Jiang, Bo, et al.
Published: (2024)
A Modular Robotic System for Autonomous Exploration and Semantic Updating in Large-Scale Indoor Environments
by: Allu, Sai Haneesh, et al.
Published: (2024)
by: Allu, Sai Haneesh, et al.
Published: (2024)
CapNav: Benchmarking Vision Language Models on Capability-conditioned Indoor Navigation
by: Su, Xia, et al.
Published: (2026)
by: Su, Xia, et al.
Published: (2026)
SynDroneVision: A Synthetic Dataset for Image-Based Drone Detection
by: Lenhard, Tamara R., et al.
Published: (2024)
by: Lenhard, Tamara R., et al.
Published: (2024)
Enhancing Vision-Language Navigation with Multimodal Event Knowledge from Real-World Indoor Tour Videos
by: Xu, Haoxuan, et al.
Published: (2026)
by: Xu, Haoxuan, et al.
Published: (2026)
Where to Perch in a Tree: Vision-Guidance for Tree-Grasping Drones
by: Dunnett, Alex, et al.
Published: (2026)
by: Dunnett, Alex, et al.
Published: (2026)
AutoTrust: Benchmarking Trustworthiness in Large Vision Language Models for Autonomous Driving
by: Xing, Shuo, et al.
Published: (2024)
by: Xing, Shuo, et al.
Published: (2024)
HetroD: A High-Fidelity Drone Dataset and Benchmark for Autonomous Driving in Heterogeneous Traffic
by: Chen, Yu-Hsiang, et al.
Published: (2026)
by: Chen, Yu-Hsiang, et al.
Published: (2026)
Safe-VLN: Collision Avoidance for Vision-and-Language Navigation of Autonomous Robots Operating in Continuous Environments
by: Yue, Lu, et al.
Published: (2023)
by: Yue, Lu, et al.
Published: (2023)
CageDroneRF: A Large-Scale RF Benchmark and Toolkit for Drone Perception
by: Rostami, Mohammad, et al.
Published: (2026)
by: Rostami, Mohammad, et al.
Published: (2026)
Asynchronous Large Language Model Enhanced Planner for Autonomous Driving
by: Chen, Yuan, et al.
Published: (2024)
by: Chen, Yuan, et al.
Published: (2024)
Distilling Multi-modal Large Language Models for Autonomous Driving
by: Hegde, Deepti, et al.
Published: (2025)
by: Hegde, Deepti, et al.
Published: (2025)
Rule-VLN: Bridging Perception and Compliance via Semantic Reasoning and Geometric Rectification
by: Wen, Jiawen, et al.
Published: (2026)
by: Wen, Jiawen, et al.
Published: (2026)
UKDM: Underwater keypoint detection and matching using underwater image enhancement techniques
by: Diaz-Garcia, Pedro, et al.
Published: (2025)
by: Diaz-Garcia, Pedro, et al.
Published: (2025)
C-CoT: Counterfactual Chain-of-Thought with Vision-Language Models for Safe Autonomous Driving
by: Tian, Kefei, et al.
Published: (2026)
by: Tian, Kefei, et al.
Published: (2026)
SimLingo: Vision-Only Closed-Loop Autonomous Driving with Language-Action Alignment
by: Renz, Katrin, et al.
Published: (2025)
by: Renz, Katrin, et al.
Published: (2025)
AIGeN: An Adversarial Approach for Instruction Generation in VLN
by: Rawal, Niyati, et al.
Published: (2024)
by: Rawal, Niyati, et al.
Published: (2024)
VLN-Zero: Rapid Exploration and Cache-Enabled Neurosymbolic Vision-Language Planning for Zero-Shot Transfer in Robot Navigation
by: Bhatt, Neel P., et al.
Published: (2025)
by: Bhatt, Neel P., et al.
Published: (2025)
Similar Items
-
CHIRLA: Comprehensive High-resolution Identification and Re-identification for Large-scale Analysis
by: Dominguez-Dager, Bessie, et al.
Published: (2025) -
AIDEN: Design and Pilot Study of an AI Assistant for the Visually Impaired
by: Marquez-Carpintero, Luis, et al.
Published: (2025) -
CADDI: An in-Class Activity Detection Dataset using IMU data from low-cost sensors
by: Marquez-Carpintero, Luis, et al.
Published: (2025) -
AdaVLN: Towards Visual Language Navigation in Continuous Indoor Environments with Moving Humans
by: Loh, Dillon, et al.
Published: (2024) -
AwareVLN: Reasoning with Self-awareness for Vision-Language Navigation
by: Guo, Wenxuan, et al.
Published: (2026)