Saved in:
Bibliographic Details
Main Authors: Tulbure, Mirela G., Caineta, Julio, Broich, Mark, Gaines, Mollie D., Rufin, Philippe, Thomas, Leon-Friedrich, Alemohammad, Hamed, Hemmerling, Jan, Hostert, Patrick
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2512.02055
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918227365330944
author Tulbure, Mirela G.
Caineta, Julio
Broich, Mark
Gaines, Mollie D.
Rufin, Philippe
Thomas, Leon-Friedrich
Alemohammad, Hamed
Hemmerling, Jan
Hostert, Patrick
author_facet Tulbure, Mirela G.
Caineta, Julio
Broich, Mark
Gaines, Mollie D.
Rufin, Philippe
Thomas, Leon-Friedrich
Alemohammad, Hamed
Hemmerling, Jan
Hostert, Patrick
contents Floods are among the most damaging weather-related hazards, and in 2024, the warmest year on record, extreme flood events affected communities across five continents. Earth observation (EO) satellites provide critical, frequent coverage for mapping inundation, yet operational accuracy depends heavily on labeled datasets and model generalization. Recent Geospatial Foundation Models (GFMs), such as ESA-IBM's TerraMind, offer improved generalizability through large-scale self-supervised pretraining, but their performance on diverse global flood events remains poorly understood. We fine-tune TerraMind for flood extent mapping using FloodsNet, a harmonized multimodal dataset containing co-located Sentinel-1 (Synthetic Aperture Radar, SAR data) and Sentinel-2 (optical) imagery for 85 flood events worldwide. We tested four configurations (base vs. large models; frozen vs. unfrozen backbones) and compared against the TerraMind Sen1Floods11 example and a U-Net trained on both FloodsNet and Sen1Floods11. The base-unfrozen configuration provided the best balance of accuracy, precision, and recall at substantially lower computational cost than the large model. The large unfrozen model achieved the highest recall. Models trained on FloodsNet outperformed the Sen1Floods11-trained example in recall with similar overall accuracy. U-Net achieved higher recall than all GFM configurations, though with slightly lower accuracy and precision. Our results demonstrate that integrating multimodal optical and SAR data and fine-tuning a GFM can enhance near-real-time flood mapping. This study provides one of the first global-scale evaluations of a GFM for flood segmentation, highlighting both its potential and current limitations for climate adaptation and disaster resilience.
format Preprint
id arxiv_https___arxiv_org_abs_2512_02055
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Leveraging AI multimodal geospatial foundation models for improved near-real-time flood mapping at a global scale
Tulbure, Mirela G.
Caineta, Julio
Broich, Mark
Gaines, Mollie D.
Rufin, Philippe
Thomas, Leon-Friedrich
Alemohammad, Hamed
Hemmerling, Jan
Hostert, Patrick
Computer Vision and Pattern Recognition
Artificial Intelligence
Floods are among the most damaging weather-related hazards, and in 2024, the warmest year on record, extreme flood events affected communities across five continents. Earth observation (EO) satellites provide critical, frequent coverage for mapping inundation, yet operational accuracy depends heavily on labeled datasets and model generalization. Recent Geospatial Foundation Models (GFMs), such as ESA-IBM's TerraMind, offer improved generalizability through large-scale self-supervised pretraining, but their performance on diverse global flood events remains poorly understood. We fine-tune TerraMind for flood extent mapping using FloodsNet, a harmonized multimodal dataset containing co-located Sentinel-1 (Synthetic Aperture Radar, SAR data) and Sentinel-2 (optical) imagery for 85 flood events worldwide. We tested four configurations (base vs. large models; frozen vs. unfrozen backbones) and compared against the TerraMind Sen1Floods11 example and a U-Net trained on both FloodsNet and Sen1Floods11. The base-unfrozen configuration provided the best balance of accuracy, precision, and recall at substantially lower computational cost than the large model. The large unfrozen model achieved the highest recall. Models trained on FloodsNet outperformed the Sen1Floods11-trained example in recall with similar overall accuracy. U-Net achieved higher recall than all GFM configurations, though with slightly lower accuracy and precision. Our results demonstrate that integrating multimodal optical and SAR data and fine-tuning a GFM can enhance near-real-time flood mapping. This study provides one of the first global-scale evaluations of a GFM for flood segmentation, highlighting both its potential and current limitations for climate adaptation and disaster resilience.
title Leveraging AI multimodal geospatial foundation models for improved near-real-time flood mapping at a global scale
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2512.02055