Can Generative Geospatial Diffusion Models Excel as Discriminative Geospatial Foundation Models?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jia, Yuru, Marsocci, Valerio, Gong, Ziyang, Yang, Xue, Vergauwen, Maarten, Nascetti, Andrea
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918139565965312
author Jia, Yuru
Marsocci, Valerio
Gong, Ziyang
Yang, Xue
Vergauwen, Maarten
Nascetti, Andrea
author_facet Jia, Yuru
Marsocci, Valerio
Gong, Ziyang
Yang, Xue
Vergauwen, Maarten
Nascetti, Andrea
contents Self-supervised learning (SSL) has revolutionized representation learning in Remote Sensing (RS), advancing Geospatial Foundation Models (GFMs) to leverage vast unlabeled satellite imagery for diverse downstream tasks. Currently, GFMs primarily employ objectives like contrastive learning or masked image modeling, owing to their proven success in learning transferable representations. However, generative diffusion models, which demonstrate the potential to capture multi-grained semantics essential for RS tasks during image generation, remain underexplored for discriminative applications. This prompts the question: can generative diffusion models also excel and serve as GFMs with sufficient discriminative power? In this work, we answer this question with SatDiFuser, a framework that transforms a diffusion-based generative geospatial foundation model into a powerful pretraining tool for discriminative RS. By systematically analyzing multi-stage, noise-dependent diffusion features, we develop three fusion strategies to effectively leverage these diverse representations. Extensive experiments on remote sensing benchmarks show that SatDiFuser outperforms state-of-the-art GFMs, achieving gains of up to +5.7% mIoU in semantic segmentation and +7.9% F1-score in classification, demonstrating the capacity of diffusion-based generative foundation models to rival or exceed discriminative GFMs. The source code is available at: https://github.com/yurujaja/SatDiFuser.
format Preprint
id arxiv_https___arxiv_org_abs_2503_07890
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Can Generative Geospatial Diffusion Models Excel as Discriminative Geospatial Foundation Models?
Jia, Yuru
Marsocci, Valerio
Gong, Ziyang
Yang, Xue
Vergauwen, Maarten
Nascetti, Andrea
Computer Vision and Pattern Recognition
Self-supervised learning (SSL) has revolutionized representation learning in Remote Sensing (RS), advancing Geospatial Foundation Models (GFMs) to leverage vast unlabeled satellite imagery for diverse downstream tasks. Currently, GFMs primarily employ objectives like contrastive learning or masked image modeling, owing to their proven success in learning transferable representations. However, generative diffusion models, which demonstrate the potential to capture multi-grained semantics essential for RS tasks during image generation, remain underexplored for discriminative applications. This prompts the question: can generative diffusion models also excel and serve as GFMs with sufficient discriminative power? In this work, we answer this question with SatDiFuser, a framework that transforms a diffusion-based generative geospatial foundation model into a powerful pretraining tool for discriminative RS. By systematically analyzing multi-stage, noise-dependent diffusion features, we develop three fusion strategies to effectively leverage these diverse representations. Extensive experiments on remote sensing benchmarks show that SatDiFuser outperforms state-of-the-art GFMs, achieving gains of up to +5.7% mIoU in semantic segmentation and +7.9% F1-score in classification, demonstrating the capacity of diffusion-based generative foundation models to rival or exceed discriminative GFMs. The source code is available at: https://github.com/yurujaja/SatDiFuser.
title Can Generative Geospatial Diffusion Models Excel as Discriminative Geospatial Foundation Models?
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.07890