GEO-Bench-2: From Performance to Capability, Rethinking Evaluation in Geospatial AI

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Simumba, Naomi, Lehmann, Nils, Fraccaro, Paolo, Alemohammad, Hamed, De Mel, Geeth, Khan, Salman, Maskey, Manil, Longepe, Nicolas, Zhu, Xiao Xiang, Kerner, Hannah, Bernabe-Moreno, Juan, Lacoste, Alexandre
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908804258463744
author Simumba, Naomi
Lehmann, Nils
Fraccaro, Paolo
Alemohammad, Hamed
De Mel, Geeth
Khan, Salman
Maskey, Manil
Longepe, Nicolas
Zhu, Xiao Xiang
Kerner, Hannah
Bernabe-Moreno, Juan
Lacoste, Alexandre
author_facet Simumba, Naomi
Lehmann, Nils
Fraccaro, Paolo
Alemohammad, Hamed
De Mel, Geeth
Khan, Salman
Maskey, Manil
Longepe, Nicolas
Zhu, Xiao Xiang
Kerner, Hannah
Bernabe-Moreno, Juan
Lacoste, Alexandre
contents Geospatial Foundation Models (GeoFMs) are transforming Earth Observation (EO), but evaluation lacks standardized protocols. GEO-Bench-2 addresses this with a comprehensive framework spanning classification, segmentation, regression, object detection, and instance segmentation across 19 permissively-licensed datasets. We introduce ''capability'' groups to rank models on datasets that share common characteristics (e.g., resolution, bands, temporality). This enables users to identify which models excel in each capability and determine which areas need improvement in future work. To support both fair comparison and methodological innovation, we define a prescriptive yet flexible evaluation protocol. This not only ensures consistency in benchmarking but also facilitates research into model adaptation strategies, a key and open challenge in advancing GeoFMs for downstream tasks. Our experiments show that no single model dominates across all tasks, confirming the specificity of the choices made during architecture design and pretraining. While models pretrained on natural images (ConvNext ImageNet, DINO V3) excel on high-resolution tasks, EO-specific models (TerraMind, Prithvi, and Clay) outperform them on multispectral applications such as agriculture and disaster response. These findings demonstrate that optimal model choice depends on task requirements, data modalities, and constraints. This shows that the goal of a single GeoFM model that performs well across all tasks remains open for future research. GEO-Bench-2 enables informed, reproducible GeoFM evaluation tailored to specific use cases. Code, data, and leaderboard for GEO-Bench-2 are publicly released under a permissive license.
format Preprint
id arxiv_https___arxiv_org_abs_2511_15658
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle GEO-Bench-2: From Performance to Capability, Rethinking Evaluation in Geospatial AI
Simumba, Naomi
Lehmann, Nils
Fraccaro, Paolo
Alemohammad, Hamed
De Mel, Geeth
Khan, Salman
Maskey, Manil
Longepe, Nicolas
Zhu, Xiao Xiang
Kerner, Hannah
Bernabe-Moreno, Juan
Lacoste, Alexandre
Computer Vision and Pattern Recognition
Artificial Intelligence
Geospatial Foundation Models (GeoFMs) are transforming Earth Observation (EO), but evaluation lacks standardized protocols. GEO-Bench-2 addresses this with a comprehensive framework spanning classification, segmentation, regression, object detection, and instance segmentation across 19 permissively-licensed datasets. We introduce ''capability'' groups to rank models on datasets that share common characteristics (e.g., resolution, bands, temporality). This enables users to identify which models excel in each capability and determine which areas need improvement in future work. To support both fair comparison and methodological innovation, we define a prescriptive yet flexible evaluation protocol. This not only ensures consistency in benchmarking but also facilitates research into model adaptation strategies, a key and open challenge in advancing GeoFMs for downstream tasks. Our experiments show that no single model dominates across all tasks, confirming the specificity of the choices made during architecture design and pretraining. While models pretrained on natural images (ConvNext ImageNet, DINO V3) excel on high-resolution tasks, EO-specific models (TerraMind, Prithvi, and Clay) outperform them on multispectral applications such as agriculture and disaster response. These findings demonstrate that optimal model choice depends on task requirements, data modalities, and constraints. This shows that the goal of a single GeoFM model that performs well across all tasks remains open for future research. GEO-Bench-2 enables informed, reproducible GeoFM evaluation tailored to specific use cases. Code, data, and leaderboard for GEO-Bench-2 are publicly released under a permissive license.
title GEO-Bench-2: From Performance to Capability, Rethinking Evaluation in Geospatial AI
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2511.15658