A Vision Centric Remote Sensing Benchmark

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Adejumo, Abduljaleel, Yeganli, Faegheh, Broni-bediako, Clifford, Xiao, Aoran, Yokoya, Naoto, Siam, Mennatullah
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918015481675776
author Adejumo, Abduljaleel
Yeganli, Faegheh
Broni-bediako, Clifford
Xiao, Aoran
Yokoya, Naoto
Siam, Mennatullah
author_facet Adejumo, Abduljaleel
Yeganli, Faegheh
Broni-bediako, Clifford
Xiao, Aoran
Yokoya, Naoto
Siam, Mennatullah
contents Multimodal Large Language Models (MLLMs) have achieved remarkable success in vision-language tasks but their remote sensing (RS) counterpart are relatively under explored. Unlike natural images, RS imagery presents unique challenges that current MLLMs struggle to handle, particularly in visual grounding and spatial reasoning. This study investigates the limitations of CLIP-based MLLMs in RS, highlighting their failure to differentiate visually distinct yet semantically similar RS images. To address this, we introduce a remote sensing multimodal visual patterns (RSMMVP) benchmark. It is designed to evaluate MLLMs in RS tasks by identifying the CLIP-blind pairs, where CLIP-based models incorrectly assign high similarity scores to visually distinct RS images. Through a visual question answering (VQA) evaluation, we analyze the performance of state-of-the-art MLLMs, revealing significant limitations in RS specific representation learning. The results provide valuable insights into the weaknesses of CLIP-based visual encoding and offer a foundation for future research to develop more effective MLLMs tailored for remote sensing applications.
format Preprint
id arxiv_https___arxiv_org_abs_2503_15816
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Vision Centric Remote Sensing Benchmark
Adejumo, Abduljaleel
Yeganli, Faegheh
Broni-bediako, Clifford
Xiao, Aoran
Yokoya, Naoto
Siam, Mennatullah
Computer Vision and Pattern Recognition
F.2.2; I.2.7
Multimodal Large Language Models (MLLMs) have achieved remarkable success in vision-language tasks but their remote sensing (RS) counterpart are relatively under explored. Unlike natural images, RS imagery presents unique challenges that current MLLMs struggle to handle, particularly in visual grounding and spatial reasoning. This study investigates the limitations of CLIP-based MLLMs in RS, highlighting their failure to differentiate visually distinct yet semantically similar RS images. To address this, we introduce a remote sensing multimodal visual patterns (RSMMVP) benchmark. It is designed to evaluate MLLMs in RS tasks by identifying the CLIP-blind pairs, where CLIP-based models incorrectly assign high similarity scores to visually distinct RS images. Through a visual question answering (VQA) evaluation, we analyze the performance of state-of-the-art MLLMs, revealing significant limitations in RS specific representation learning. The results provide valuable insights into the weaknesses of CLIP-based visual encoding and offer a foundation for future research to develop more effective MLLMs tailored for remote sensing applications.
title A Vision Centric Remote Sensing Benchmark
topic Computer Vision and Pattern Recognition
F.2.2; I.2.7
url https://arxiv.org/abs/2503.15816