Pretrain Where? Investigating How Pretraining Data Diversity Impacts Geospatial Foundation Model Performance

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kaur, Amandeep, Purohit, Mirali, Muhawenayo, Gedeon, Rolf, Esther, Kerner, Hannah
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911616583335936
author Kaur, Amandeep
Purohit, Mirali
Muhawenayo, Gedeon
Rolf, Esther
Kerner, Hannah
author_facet Kaur, Amandeep
Purohit, Mirali
Muhawenayo, Gedeon
Rolf, Esther
Kerner, Hannah
contents New geospatial foundation models introduce a new model architecture and pretraining dataset, often sampled using different notions of data diversity. Performance differences are largely attributed to the model architecture or input modalities, while the role of the pretraining dataset is rarely studied. To address this research gap, we conducted a systematic study on how the geographic composition of pretraining data affects a model's downstream performance. We created global and per-continent pretraining datasets and evaluated them on global and per-continent downstream datasets. We found that the pretraining dataset from Europe outperformed global and continent-specific pretraining datasets on both global and local downstream evaluations. To investigate the factors influencing a pretraining dataset's downstream performance, we analysed 10 pretraining datasets using diversity across continents, biomes, landcover and spectral values. We found that only spectral diversity was strongly correlated with performance, while others were weakly correlated. This finding establishes a new dimension of diversity to be accounted for when creating a high-performing pretraining dataset. We open-sourced 7 new pretraining datasets, pretrained models, and our experimental framework at https://github.com/kerner-lab/pretrain-where.
format Preprint
id arxiv_https___arxiv_org_abs_2604_21104
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Pretrain Where? Investigating How Pretraining Data Diversity Impacts Geospatial Foundation Model Performance
Kaur, Amandeep
Purohit, Mirali
Muhawenayo, Gedeon
Rolf, Esther
Kerner, Hannah
Computer Vision and Pattern Recognition
Machine Learning
New geospatial foundation models introduce a new model architecture and pretraining dataset, often sampled using different notions of data diversity. Performance differences are largely attributed to the model architecture or input modalities, while the role of the pretraining dataset is rarely studied. To address this research gap, we conducted a systematic study on how the geographic composition of pretraining data affects a model's downstream performance. We created global and per-continent pretraining datasets and evaluated them on global and per-continent downstream datasets. We found that the pretraining dataset from Europe outperformed global and continent-specific pretraining datasets on both global and local downstream evaluations. To investigate the factors influencing a pretraining dataset's downstream performance, we analysed 10 pretraining datasets using diversity across continents, biomes, landcover and spectral values. We found that only spectral diversity was strongly correlated with performance, while others were weakly correlated. This finding establishes a new dimension of diversity to be accounted for when creating a high-performing pretraining dataset. We open-sourced 7 new pretraining datasets, pretrained models, and our experimental framework at https://github.com/kerner-lab/pretrain-where.
title Pretrain Where? Investigating How Pretraining Data Diversity Impacts Geospatial Foundation Model Performance
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2604.21104