UAVBench and UAVIT-1M: Benchmarking and Enhancing MLLMs for Low-Altitude UAV Vision-Language Understanding
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhan, Yang, Yuan, Yuan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
MM-UAVBench: How Well Do Multimodal Large Language Models See, Think, and Plan in Low-Altitude UAV Scenarios?
di: Dai, Shiqi, et al.
Pubblicazione: (2025)
di: Dai, Shiqi, et al.
Pubblicazione: (2025)
Law of Vision Representation in MLLMs
di: Yang, Shijia, et al.
Pubblicazione: (2024)
di: Yang, Shijia, et al.
Pubblicazione: (2024)
From Recognition to Reasoning: Benchmarking and Enhancing MLLMs on Real-World Receipt Document Understanding
di: Wang, Yandi, et al.
Pubblicazione: (2026)
di: Wang, Yandi, et al.
Pubblicazione: (2026)
EVD4UAV: An Altitude-Sensitive Benchmark to Evade Vehicle Detection in UAV
di: Sun, Huiming, et al.
Pubblicazione: (2024)
di: Sun, Huiming, et al.
Pubblicazione: (2024)
Do MLLMs Understand Pointing? Benchmarking and Enhancing Referential Reasoning in Egocentric Vision
di: Li, Chentao, et al.
Pubblicazione: (2026)
di: Li, Chentao, et al.
Pubblicazione: (2026)
Prototype-Based Low Altitude UAV Semantic Segmentation
di: Zhang, Da, et al.
Pubblicazione: (2026)
di: Zhang, Da, et al.
Pubblicazione: (2026)
SkyEyeGPT: Unifying Remote Sensing Vision-Language Tasks via Instruction Tuning with Large Language Model
di: Zhan, Yang, et al.
Pubblicazione: (2024)
di: Zhan, Yang, et al.
Pubblicazione: (2024)
Benchmarking Large and Small MLLMs
di: Feng, Xuelu, et al.
Pubblicazione: (2025)
di: Feng, Xuelu, et al.
Pubblicazione: (2025)
GOBench: Benchmarking Geometric Optics Generation and Understanding of MLLMs
di: Zhu, Xiaorong, et al.
Pubblicazione: (2025)
di: Zhu, Xiaorong, et al.
Pubblicazione: (2025)
SARLANG-1M: A Benchmark for Vision-Language Modeling in SAR Image Understanding
di: Wei, Yimin, et al.
Pubblicazione: (2025)
di: Wei, Yimin, et al.
Pubblicazione: (2025)
VISTA: Enhancing Vision-Text Alignment in MLLMs via Cross-Modal Mutual Information Maximization
di: Li, Mingxiao, et al.
Pubblicazione: (2025)
di: Li, Mingxiao, et al.
Pubblicazione: (2025)
Leveraging Large Vision Model for Multi-UAV Co-perception in Low-Altitude Wireless Networks
di: Xu, Yunting, et al.
Pubblicazione: (2026)
di: Xu, Yunting, et al.
Pubblicazione: (2026)
UAV-Borne Mapping Algorithms for Low-Altitude and High-Speed Drone Applications
di: Zhang, Jincheng, et al.
Pubblicazione: (2024)
di: Zhang, Jincheng, et al.
Pubblicazione: (2024)
Human Cognitive Benchmarks Reveal Foundational Visual Gaps in MLLMs
di: Huang, Jen-Tse, et al.
Pubblicazione: (2025)
di: Huang, Jen-Tse, et al.
Pubblicazione: (2025)
E-VAds: An E-commerce Short Videos Understanding Benchmark for MLLMs
di: Liu, Xianjie, et al.
Pubblicazione: (2026)
di: Liu, Xianjie, et al.
Pubblicazione: (2026)
Towards Realistic UAV Vision-Language Navigation: Platform, Benchmark, and Methodology
di: Wang, Xiangyu, et al.
Pubblicazione: (2024)
di: Wang, Xiangyu, et al.
Pubblicazione: (2024)
IF-Bench: Benchmarking and Enhancing MLLMs for Infrared Images with Generative Visual Prompting
di: Zhang, Tao, et al.
Pubblicazione: (2025)
di: Zhang, Tao, et al.
Pubblicazione: (2025)
AdaptVision: Dynamic Input Scaling in MLLMs for Versatile Scene Understanding
di: Wang, Yonghui, et al.
Pubblicazione: (2024)
di: Wang, Yonghui, et al.
Pubblicazione: (2024)
Math Blind: Failures in Diagram Understanding Undermine Reasoning in MLLMs
di: Sun, Yanpeng, et al.
Pubblicazione: (2025)
di: Sun, Yanpeng, et al.
Pubblicazione: (2025)
Focus on What Really Matters in Low-Altitude Governance: A Management-Centric Multi-Modal Benchmark with Implicitly Coordinated Vision-Language Reasoning Framework
di: Chang, Hao, et al.
Pubblicazione: (2026)
di: Chang, Hao, et al.
Pubblicazione: (2026)
Can MLLMs Reason Beyond Language? VisReason: A Comprehensive Benchmark for Vision-Centric Reasoning
di: Guo, Longteng, et al.
Pubblicazione: (2026)
di: Guo, Longteng, et al.
Pubblicazione: (2026)
PhysToolBench: Benchmarking Physical Tool Understanding for MLLMs
di: Zhang, Zixin, et al.
Pubblicazione: (2025)
di: Zhang, Zixin, et al.
Pubblicazione: (2025)
MathFlow: Enhancing the Perceptual Flow of MLLMs for Visual Mathematical Problems
di: Chen, Shuhang, et al.
Pubblicazione: (2025)
di: Chen, Shuhang, et al.
Pubblicazione: (2025)
Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking
di: Chung, Sangyun, et al.
Pubblicazione: (2024)
di: Chung, Sangyun, et al.
Pubblicazione: (2024)
SpatialBot: Precise Spatial Understanding with Vision Language Models
di: Cai, Wenxiao, et al.
Pubblicazione: (2024)
di: Cai, Wenxiao, et al.
Pubblicazione: (2024)
Benchmarking the Robustness of UAV Tracking Against Common Corruptions
di: Liu, Xiaoqiong, et al.
Pubblicazione: (2024)
di: Liu, Xiaoqiong, et al.
Pubblicazione: (2024)
TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning
di: Wu, Tao, et al.
Pubblicazione: (2025)
di: Wu, Tao, et al.
Pubblicazione: (2025)
Open Eyes, Then Reason: Fine-grained Visual Mathematical Understanding in MLLMs
di: Zhang, Shan, et al.
Pubblicazione: (2025)
di: Zhang, Shan, et al.
Pubblicazione: (2025)
Touch-R1: Reinforcing Touch Reasoning in MLLMs
di: Lai, Yingxin, et al.
Pubblicazione: (2026)
di: Lai, Yingxin, et al.
Pubblicazione: (2026)
VLDrive: Vision-Augmented Lightweight MLLMs for Efficient Language-grounded Autonomous Driving
di: Zhang, Ruifei, et al.
Pubblicazione: (2025)
di: Zhang, Ruifei, et al.
Pubblicazione: (2025)
TUBench: Benchmarking Large Vision-Language Models on Trustworthiness with Unanswerable Questions
di: He, Xingwei, et al.
Pubblicazione: (2024)
di: He, Xingwei, et al.
Pubblicazione: (2024)
Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark
di: Hao, Yunzhuo, et al.
Pubblicazione: (2025)
di: Hao, Yunzhuo, et al.
Pubblicazione: (2025)
HUGE-Bench: A Benchmark for High-Level UAV Vision-Language-Action Tasks
di: Guo, Jingyu, et al.
Pubblicazione: (2026)
di: Guo, Jingyu, et al.
Pubblicazione: (2026)
AdaptFly: Prompt-Guided Adaptation of Foundation Models for Low-Altitude UAV Networks
di: Chen, Jiao, et al.
Pubblicazione: (2025)
di: Chen, Jiao, et al.
Pubblicazione: (2025)
Where Does Vision Meet Language? Understanding and Refining Visual Fusion in MLLMs via Contrastive Attention
di: Song, Shezheng, et al.
Pubblicazione: (2026)
di: Song, Shezheng, et al.
Pubblicazione: (2026)
RIS-LAD: A Benchmark and Model for Referring Low-Altitude Drone Image Segmentation
di: Ye, Kai, et al.
Pubblicazione: (2025)
di: Ye, Kai, et al.
Pubblicazione: (2025)
Patch-as-Decodable-Token: Towards Unified Multi-Modal Vision Tasks in MLLMs
di: Su, Yongyi, et al.
Pubblicazione: (2025)
di: Su, Yongyi, et al.
Pubblicazione: (2025)
RT-DETR++ for UAV Object Detection
di: Shufang, Yuan
Pubblicazione: (2025)
di: Shufang, Yuan
Pubblicazione: (2025)
UWBench: A Comprehensive Vision-Language Benchmark for Underwater Understanding
di: Zhang, Da, et al.
Pubblicazione: (2025)
di: Zhang, Da, et al.
Pubblicazione: (2025)
Universal Skeleton Understanding via Differentiable Rendering and MLLMs
di: Wang, Ziyi, et al.
Pubblicazione: (2026)
di: Wang, Ziyi, et al.
Pubblicazione: (2026)
Documenti analoghi
-
MM-UAVBench: How Well Do Multimodal Large Language Models See, Think, and Plan in Low-Altitude UAV Scenarios?
di: Dai, Shiqi, et al.
Pubblicazione: (2025) -
Law of Vision Representation in MLLMs
di: Yang, Shijia, et al.
Pubblicazione: (2024) -
From Recognition to Reasoning: Benchmarking and Enhancing MLLMs on Real-World Receipt Document Understanding
di: Wang, Yandi, et al.
Pubblicazione: (2026) -
EVD4UAV: An Altitude-Sensitive Benchmark to Evade Vehicle Detection in UAV
di: Sun, Huiming, et al.
Pubblicazione: (2024) -
Do MLLMs Understand Pointing? Benchmarking and Enhancing Referential Reasoning in Egocentric Vision
di: Li, Chentao, et al.
Pubblicazione: (2026)