TerraFM: A Scalable Foundation Model for Unified Multisensor Earth Observation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Danish, Muhammad Sohail, Munir, Muhammad Akhtar, Shah, Syed Roshaan Ali, Khan, Muhammad Haris, Anwer, Rao Muhammad, Laaksonen, Jorma, Khan, Fahad Shahbaz, Khan, Salman
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910992505503744
author Danish, Muhammad Sohail
Munir, Muhammad Akhtar
Shah, Syed Roshaan Ali
Khan, Muhammad Haris
Anwer, Rao Muhammad
Laaksonen, Jorma
Khan, Fahad Shahbaz
Khan, Salman
author_facet Danish, Muhammad Sohail
Munir, Muhammad Akhtar
Shah, Syed Roshaan Ali
Khan, Muhammad Haris
Anwer, Rao Muhammad
Laaksonen, Jorma
Khan, Fahad Shahbaz
Khan, Salman
contents Modern Earth observation (EO) increasingly leverages deep learning to harness the scale and diversity of satellite imagery across sensors and regions. While recent foundation models have demonstrated promising generalization across EO tasks, many remain limited by the scale, geographical coverage, and spectral diversity of their training data, factors critical for learning globally transferable representations. In this work, we introduce TerraFM, a scalable self-supervised learning model that leverages globally distributed Sentinel-1 and Sentinel-2 imagery, combined with large spatial tiles and land-cover aware sampling to enrich spatial and semantic coverage. By treating sensing modalities as natural augmentations in our self-supervised approach, we unify radar and optical inputs via modality-specific patch embeddings and adaptive cross-attention fusion. Our training strategy integrates local-global contrastive learning and introduces a dual-centering mechanism that incorporates class-frequency-aware regularization to address long-tailed distributions in land cover.TerraFM achieves strong generalization on both classification and segmentation tasks, outperforming prior models on GEO-Bench and Copernicus-Bench. Our code and pretrained models are publicly available at: https://github.com/mbzuai-oryx/TerraFM .
format Preprint
id arxiv_https___arxiv_org_abs_2506_06281
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle TerraFM: A Scalable Foundation Model for Unified Multisensor Earth Observation
Danish, Muhammad Sohail
Munir, Muhammad Akhtar
Shah, Syed Roshaan Ali
Khan, Muhammad Haris
Anwer, Rao Muhammad
Laaksonen, Jorma
Khan, Fahad Shahbaz
Khan, Salman
Computer Vision and Pattern Recognition
Modern Earth observation (EO) increasingly leverages deep learning to harness the scale and diversity of satellite imagery across sensors and regions. While recent foundation models have demonstrated promising generalization across EO tasks, many remain limited by the scale, geographical coverage, and spectral diversity of their training data, factors critical for learning globally transferable representations. In this work, we introduce TerraFM, a scalable self-supervised learning model that leverages globally distributed Sentinel-1 and Sentinel-2 imagery, combined with large spatial tiles and land-cover aware sampling to enrich spatial and semantic coverage. By treating sensing modalities as natural augmentations in our self-supervised approach, we unify radar and optical inputs via modality-specific patch embeddings and adaptive cross-attention fusion. Our training strategy integrates local-global contrastive learning and introduces a dual-centering mechanism that incorporates class-frequency-aware regularization to address long-tailed distributions in land cover.TerraFM achieves strong generalization on both classification and segmentation tasks, outperforming prior models on GEO-Bench and Copernicus-Bench. Our code and pretrained models are publicly available at: https://github.com/mbzuai-oryx/TerraFM .
title TerraFM: A Scalable Foundation Model for Unified Multisensor Earth Observation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.06281