EarthMind: Leveraging Cross-Sensor Data for Advanced Earth Observation Interpretation with a Unified Multimodal LLM

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shu, Yan, Ren, Bin, Xiong, Zhitong, Paudel, Danda Pani, Van Gool, Luc, Demir, Begüm, Sebe, Nicu, Rota, Paolo
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912610830516224
author Shu, Yan
Ren, Bin
Xiong, Zhitong
Paudel, Danda Pani
Van Gool, Luc
Demir, Begüm
Sebe, Nicu
Rota, Paolo
author_facet Shu, Yan
Ren, Bin
Xiong, Zhitong
Paudel, Danda Pani
Van Gool, Luc
Demir, Begüm
Sebe, Nicu
Rota, Paolo
contents Earth Observation (EO) data analysis is vital for monitoring environmental and human dynamics. Recent Multimodal Large Language Models (MLLMs) show potential in EO understanding but remain restricted to single-sensor inputs, overlooking the complementarity across heterogeneous modalities. We propose EarthMind, a unified vision-language framework that handles both single- and cross-sensor inputs via an innovative hierarchical cross-modal attention (ie, HCA) design. Specifically, HCA hierarchically captures visual relationships across sensors and aligns them with language queries, enabling adaptive fusion of optical and Synthetic Aperture Radar (SAR) features. To support cross-sensor learning, we curate FusionEO, a 30K-pair dataset with diverse annotations, and establish EarthMind-Bench, a 2,841-pair benchmark with expert annotations for perception and reasoning tasks. Extensive experiments show that EarthMind achieves state-of-the-art results on EarthMind-Bench and surpasses existing MLLMs on multiple EO benchmarks.
format Preprint
id arxiv_https___arxiv_org_abs_2506_01667
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle EarthMind: Leveraging Cross-Sensor Data for Advanced Earth Observation Interpretation with a Unified Multimodal LLM
Shu, Yan
Ren, Bin
Xiong, Zhitong
Paudel, Danda Pani
Van Gool, Luc
Demir, Begüm
Sebe, Nicu
Rota, Paolo
Computer Vision and Pattern Recognition
Earth Observation (EO) data analysis is vital for monitoring environmental and human dynamics. Recent Multimodal Large Language Models (MLLMs) show potential in EO understanding but remain restricted to single-sensor inputs, overlooking the complementarity across heterogeneous modalities. We propose EarthMind, a unified vision-language framework that handles both single- and cross-sensor inputs via an innovative hierarchical cross-modal attention (ie, HCA) design. Specifically, HCA hierarchically captures visual relationships across sensors and aligns them with language queries, enabling adaptive fusion of optical and Synthetic Aperture Radar (SAR) features. To support cross-sensor learning, we curate FusionEO, a 30K-pair dataset with diverse annotations, and establish EarthMind-Bench, a 2,841-pair benchmark with expert annotations for perception and reasoning tasks. Extensive experiments show that EarthMind achieves state-of-the-art results on EarthMind-Bench and surpasses existing MLLMs on multiple EO benchmarks.
title EarthMind: Leveraging Cross-Sensor Data for Advanced Earth Observation Interpretation with a Unified Multimodal LLM
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.01667