Is your VLM Sky-Ready? A Comprehensive Spatial Intelligence Benchmark for UAV Navigation
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Lingfeng, Zhang, Yuchen, Li, Hongsheng, Fu, Haoxiang, Tang, Yingbo, Ye, Hangjun, Chen, Long, Liang, Xiaojun, Hao, Xiaoshuai, Ding, Wenbo |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SocialNav-Map: Dynamic Mapping with Human Trajectory Prediction for Zero-Shot Social Navigation
by: Zhang, Lingfeng, et al.
Published: (2025)
by: Zhang, Lingfeng, et al.
Published: (2025)
RoboAfford++: A Generative AI-Enhanced Dataset for Multimodal Affordance Learning in Robotic Manipulation and Navigation
by: Hao, Xiaoshuai, et al.
Published: (2025)
by: Hao, Xiaoshuai, et al.
Published: (2025)
Team Xiaomi EV-AD VLA: Caption-Guided Retrieval System for Cross-Modal Drone Navigation -- Technical Report for IROS 2025 RoboSense Challenge Track 4
by: Zhang, Lingfeng, et al.
Published: (2025)
by: Zhang, Lingfeng, et al.
Published: (2025)
Learning to Navigate Socially Through Proactive Risk Perception
by: Xiao, Erjia, et al.
Published: (2025)
by: Xiao, Erjia, et al.
Published: (2025)
$NavA^3$: Understanding Any Instruction, Navigating Anywhere, Finding Anything
by: Zhang, Lingfeng, et al.
Published: (2025)
by: Zhang, Lingfeng, et al.
Published: (2025)
Walk With Me: Long-Horizon Social Navigation for Human-Centric Outdoor Assistance
by: Zhang, Lingfeng, et al.
Published: (2026)
by: Zhang, Lingfeng, et al.
Published: (2026)
Is Your VLM for Autonomous Driving Safety-Ready? A Comprehensive Benchmark for Evaluating External and In-Cabin Risks
by: Meng, Xianhui, et al.
Published: (2025)
by: Meng, Xianhui, et al.
Published: (2025)
Thinking in Text and Images: Interleaved Vision--Language Reasoning Traces for Long-Horizon Robot Manipulation
by: Liu, Jinkun, et al.
Published: (2026)
by: Liu, Jinkun, et al.
Published: (2026)
SEF-MAP: Subspace-Decomposed Expert Fusion for Robust Multimodal HD Map Prediction
by: Fu, Haoxiang, et al.
Published: (2026)
by: Fu, Haoxiang, et al.
Published: (2026)
Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought
by: Zhang, Shuyi, et al.
Published: (2025)
by: Zhang, Shuyi, et al.
Published: (2025)
Weather-Conditioned Branch Routing for Robust LiDAR-Radar 3D Object Detection
by: Li, Hongsheng, et al.
Published: (2026)
by: Li, Hongsheng, et al.
Published: (2026)
VLM-RRT: Vision Language Model Guided RRT Search for Autonomous UAV Navigation
by: Ye, Jianlin, et al.
Published: (2025)
by: Ye, Jianlin, et al.
Published: (2025)
Securing the Skies: A Comprehensive Survey on Anti-UAV Methods, Benchmarking, and Future Directions
by: Dong, Yifei, et al.
Published: (2025)
by: Dong, Yifei, et al.
Published: (2025)
Towards Realistic UAV Vision-Language Navigation: Platform, Benchmark, and Methodology
by: Wang, Xiangyu, et al.
Published: (2024)
by: Wang, Xiangyu, et al.
Published: (2024)
Chapter Reverse modeling per la stampa 3D di complessi monumentali
by: Fu, Hangjun
Published: (2024)
by: Fu, Hangjun
Published: (2024)
VFLAIR-LLM: A Comprehensive Framework and Benchmark for Split Learning of LLMs
by: Gu, Zixuan, et al.
Published: (2025)
by: Gu, Zixuan, et al.
Published: (2025)
When Engineering Outruns Intelligence: Rethinking Instruction-Guided Navigation
by: Aghaei, Matin, et al.
Published: (2025)
by: Aghaei, Matin, et al.
Published: (2025)
DriveWorld-VLA: Unified Latent-Space World Modeling with Vision-Language-Action for Autonomous Driving
by: jia, Feiyang, et al.
Published: (2026)
by: jia, Feiyang, et al.
Published: (2026)
Face-Human-Bench: A Comprehensive Benchmark of Face and Human Understanding for Multi-modal Assistants
by: Qin, Lixiong, et al.
Published: (2025)
by: Qin, Lixiong, et al.
Published: (2025)
SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence
by: Zeng, Zhitao, et al.
Published: (2025)
by: Zeng, Zhitao, et al.
Published: (2025)
AffordGrasp: In-Context Affordance Reasoning for Open-Vocabulary Task-Oriented Grasping in Clutter
by: Tang, Yingbo, et al.
Published: (2025)
by: Tang, Yingbo, et al.
Published: (2025)
PET-F2I: A Comprehensive Benchmark and Parameter-Efficient Fine-Tuning of LLMs for PET/CT Report Impression Generation
by: Liu, Yuchen, et al.
Published: (2026)
by: Liu, Yuchen, et al.
Published: (2026)
SoraNav: Adaptive UAV Task-Centric Navigation via Zeroshot VLM Reasoning
by: Song, Hongyu, et al.
Published: (2025)
by: Song, Hongyu, et al.
Published: (2025)
MapNav: A Novel Memory Representation via Annotated Semantic Maps for Vision-and-Language Navigation
by: Zhang, Lingfeng, et al.
Published: (2025)
by: Zhang, Lingfeng, et al.
Published: (2025)
SpatialScore: Towards Comprehensive Evaluation for Spatial Intelligence
by: Wu, Haoning, et al.
Published: (2025)
by: Wu, Haoning, et al.
Published: (2025)
Task-Level AI Readiness Assessment for Business Process Management:The T-IPO Model and LARA Matrix in Financial-Services IT Operations
by: Li, Mingjun, et al.
Published: (2026)
by: Li, Mingjun, et al.
Published: (2026)
Are VLMs Lost Between Sky and Space? LinkS$^2$Bench for UAV-Satellite Dynamic Cross-View Spatial Intelligence
by: Liu, Dian, et al.
Published: (2026)
by: Liu, Dian, et al.
Published: (2026)
Are Video Models Ready as Zero-Shot Reasoners? An Empirical Study with the MME-CoF Benchmark
by: Guo, Ziyu, et al.
Published: (2025)
by: Guo, Ziyu, et al.
Published: (2025)
Can Vision-Language Models Think from the Sky? Unifying UAV Reasoning and Generation
by: Sun, Jintao, et al.
Published: (2026)
by: Sun, Jintao, et al.
Published: (2026)
IndoorUAV: Benchmarking Vision-Language UAV Navigation in Continuous Indoor Environments
by: Liu, Xu, et al.
Published: (2025)
by: Liu, Xu, et al.
Published: (2025)
Stepping VLMs onto the Court: Benchmarking Spatial Intelligence in Sports
by: Yang, Yuchen, et al.
Published: (2026)
by: Yang, Yuchen, et al.
Published: (2026)
Is AI Ready for Multimodal Hate Speech Detection? A Comprehensive Dataset and Benchmark Evaluation
by: Xing, Rui, et al.
Published: (2026)
by: Xing, Rui, et al.
Published: (2026)
OmniSpatial: Towards Comprehensive Spatial Reasoning Benchmark for Vision Language Models
by: Jia, Mengdi, et al.
Published: (2025)
by: Jia, Mengdi, et al.
Published: (2025)
MMSI-Video-Bench: A Holistic Benchmark for Video-Based Spatial Intelligence
by: Lin, Jingli, et al.
Published: (2025)
by: Lin, Jingli, et al.
Published: (2025)
Benchmarking and Enhancing VLM for Compressed Image Understanding
by: Zhang, Zifu, et al.
Published: (2025)
by: Zhang, Zifu, et al.
Published: (2025)
UAV-Flow Colosseo: A Real-World Benchmark for Flying-on-a-Word UAV Imitation Learning
by: Wang, Xiangyu, et al.
Published: (2025)
by: Wang, Xiangyu, et al.
Published: (2025)
ReVSI: Rebuilding Visual Spatial Intelligence Evaluation for Accurate Assessment of VLM 3D Reasoning
by: Zhang, Yiming, et al.
Published: (2026)
by: Zhang, Yiming, et al.
Published: (2026)
DSI-Bench: A Benchmark for Dynamic Spatial Intelligence
by: Zhang, Ziang, et al.
Published: (2025)
by: Zhang, Ziang, et al.
Published: (2025)
VLM-Pruner: Buffering for Spatial Sparsity in an Efficient VLM Centrifugal Token Pruning Paradigm
by: Wu, Zhenkai, et al.
Published: (2025)
by: Wu, Zhenkai, et al.
Published: (2025)
PIGEON: VLM-Driven Object Navigation via Points of Interest Selection
by: Peng, Cheng, et al.
Published: (2025)
by: Peng, Cheng, et al.
Published: (2025)
Similar Items
-
SocialNav-Map: Dynamic Mapping with Human Trajectory Prediction for Zero-Shot Social Navigation
by: Zhang, Lingfeng, et al.
Published: (2025) -
RoboAfford++: A Generative AI-Enhanced Dataset for Multimodal Affordance Learning in Robotic Manipulation and Navigation
by: Hao, Xiaoshuai, et al.
Published: (2025) -
Team Xiaomi EV-AD VLA: Caption-Guided Retrieval System for Cross-Modal Drone Navigation -- Technical Report for IROS 2025 RoboSense Challenge Track 4
by: Zhang, Lingfeng, et al.
Published: (2025) -
Learning to Navigate Socially Through Proactive Risk Perception
by: Xiao, Erjia, et al.
Published: (2025) -
$NavA^3$: Understanding Any Instruction, Navigating Anywhere, Finding Anything
by: Zhang, Lingfeng, et al.
Published: (2025)