SARLANG-1M: A Benchmark for Vision-Language Modeling in SAR Image Understanding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wei, Yimin, Xiao, Aoran, Ren, Yexian, Zhu, Yuting, Chen, Hongruixuan, Xia, Junshi, Yokoya, Naoto
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917976544903168
author Wei, Yimin
Xiao, Aoran
Ren, Yexian
Zhu, Yuting
Chen, Hongruixuan
Xia, Junshi
Yokoya, Naoto
author_facet Wei, Yimin
Xiao, Aoran
Ren, Yexian
Zhu, Yuting
Chen, Hongruixuan
Xia, Junshi
Yokoya, Naoto
contents Synthetic Aperture Radar (SAR) is a crucial remote sensing technology, enabling all-weather, day-and-night observation with strong surface penetration for precise and continuous environmental monitoring and analysis. However, SAR image interpretation remains challenging due to its complex physical imaging mechanisms and significant visual disparities from human perception. Recently, Vision-Language Models (VLMs) have demonstrated remarkable success in RGB image understanding, offering powerful open-vocabulary interpretation and flexible language interaction. However, their application to SAR images is severely constrained by the absence of SAR-specific knowledge in their training distributions, leading to suboptimal performance. To address this limitation, we introduce SARLANG-1M, a large-scale benchmark tailored for multimodal SAR image understanding, with a primary focus on integrating SAR with textual modality. SARLANG-1M comprises more than 1 million high-quality SAR image-text pairs collected from over 59 cities worldwide. It features hierarchical resolutions (ranging from 0.1 to 25 meters), fine-grained semantic descriptions (including both concise and detailed captions), diverse remote sensing categories (1,696 object types and 16 land cover classes), and multi-task question-answering pairs spanning seven applications and 1,012 question types. Extensive experiments on mainstream VLMs demonstrate that fine-tuning with SARLANG-1M significantly enhances their performance in SAR image interpretation, reaching performance comparable to human experts. The dataset and code will be made publicly available at https://github.com/Jimmyxichen/SARLANG-1M.
format Preprint
id arxiv_https___arxiv_org_abs_2504_03254
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SARLANG-1M: A Benchmark for Vision-Language Modeling in SAR Image Understanding
Wei, Yimin
Xiao, Aoran
Ren, Yexian
Zhu, Yuting
Chen, Hongruixuan
Xia, Junshi
Yokoya, Naoto
Computer Vision and Pattern Recognition
Synthetic Aperture Radar (SAR) is a crucial remote sensing technology, enabling all-weather, day-and-night observation with strong surface penetration for precise and continuous environmental monitoring and analysis. However, SAR image interpretation remains challenging due to its complex physical imaging mechanisms and significant visual disparities from human perception. Recently, Vision-Language Models (VLMs) have demonstrated remarkable success in RGB image understanding, offering powerful open-vocabulary interpretation and flexible language interaction. However, their application to SAR images is severely constrained by the absence of SAR-specific knowledge in their training distributions, leading to suboptimal performance. To address this limitation, we introduce SARLANG-1M, a large-scale benchmark tailored for multimodal SAR image understanding, with a primary focus on integrating SAR with textual modality. SARLANG-1M comprises more than 1 million high-quality SAR image-text pairs collected from over 59 cities worldwide. It features hierarchical resolutions (ranging from 0.1 to 25 meters), fine-grained semantic descriptions (including both concise and detailed captions), diverse remote sensing categories (1,696 object types and 16 land cover classes), and multi-task question-answering pairs spanning seven applications and 1,012 question types. Extensive experiments on mainstream VLMs demonstrate that fine-tuning with SARLANG-1M significantly enhances their performance in SAR image interpretation, reaching performance comparable to human experts. The dataset and code will be made publicly available at https://github.com/Jimmyxichen/SARLANG-1M.
title SARLANG-1M: A Benchmark for Vision-Language Modeling in SAR Image Understanding
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2504.03254