VLRS-Bench: A Vision-Language Reasoning Benchmark for Remote Sensing

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Luo, Zhiming, Wang, Di, Guo, Haonan, Zhang, Jing, Du, Bo
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914563953262592
author Luo, Zhiming
Wang, Di
Guo, Haonan
Zhang, Jing
Du, Bo
author_facet Luo, Zhiming
Wang, Di
Guo, Haonan
Zhang, Jing
Du, Bo
contents Recent advancements in Multimodal Large Language Models (MLLMs) have enabled complex reasoning. However, existing remote sensing (RS) benchmarks remain heavily biased toward perception tasks, such as object recognition and scene classification. This limitation hinders the development of MLLMs for cognitively demanding RS applications. To address this, we propose a Vision Language ReaSoning Benchmark (VLRS-Bench), which is the first benchmark exclusively dedicated to complex RS reasoning. Structured across the three core dimensions of Cognition, Decision, and Prediction, VLRS-Bench comprises 2,000 question-answer pairs with an average question length of 130.19 words, spanning 14 tasks and up to eight temporal phases. VLRS-Bench is constructed via a specialized pipeline that integrates RS-specific priors and expert knowledge to ensure geospatial realism and reasoning complexity. Experimental results reveal significant bottlenecks in existing state-of-the-art MLLMs, providing critical insights for advancing multimodal reasoning within the remote sensing community. The project repository is available at https://github.com/MiliLab/VLRS-Bench.
format Preprint
id arxiv_https___arxiv_org_abs_2602_07045
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle VLRS-Bench: A Vision-Language Reasoning Benchmark for Remote Sensing
Luo, Zhiming
Wang, Di
Guo, Haonan
Zhang, Jing
Du, Bo
Computer Vision and Pattern Recognition
Artificial Intelligence
Recent advancements in Multimodal Large Language Models (MLLMs) have enabled complex reasoning. However, existing remote sensing (RS) benchmarks remain heavily biased toward perception tasks, such as object recognition and scene classification. This limitation hinders the development of MLLMs for cognitively demanding RS applications. To address this, we propose a Vision Language ReaSoning Benchmark (VLRS-Bench), which is the first benchmark exclusively dedicated to complex RS reasoning. Structured across the three core dimensions of Cognition, Decision, and Prediction, VLRS-Bench comprises 2,000 question-answer pairs with an average question length of 130.19 words, spanning 14 tasks and up to eight temporal phases. VLRS-Bench is constructed via a specialized pipeline that integrates RS-specific priors and expert knowledge to ensure geospatial realism and reasoning complexity. Experimental results reveal significant bottlenecks in existing state-of-the-art MLLMs, providing critical insights for advancing multimodal reasoning within the remote sensing community. The project repository is available at https://github.com/MiliLab/VLRS-Bench.
title VLRS-Bench: A Vision-Language Reasoning Benchmark for Remote Sensing
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2602.07045