TerraMind: Large-Scale Generative Multimodality for Earth Observation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Jakubik, Johannes, Yang, Felix, Blumenstiel, Benedikt, Scheurer, Erik, Sedona, Rocco, Maurogiovanni, Stefano, Bosmans, Jente, Dionelis, Nikolaos, Marsocci, Valerio, Kopp, Niklas, Ramachandran, Rahul, Fraccaro, Paolo, Brunschwiler, Thomas, Cavallaro, Gabriele, Bernabe-Moreno, Juan, Longépé, Nicolas
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908529313447936
author Jakubik, Johannes
Yang, Felix
Blumenstiel, Benedikt
Scheurer, Erik
Sedona, Rocco
Maurogiovanni, Stefano
Bosmans, Jente
Dionelis, Nikolaos
Marsocci, Valerio
Kopp, Niklas
Ramachandran, Rahul
Fraccaro, Paolo
Brunschwiler, Thomas
Cavallaro, Gabriele
Bernabe-Moreno, Juan
Longépé, Nicolas
author_facet Jakubik, Johannes
Yang, Felix
Blumenstiel, Benedikt
Scheurer, Erik
Sedona, Rocco
Maurogiovanni, Stefano
Bosmans, Jente
Dionelis, Nikolaos
Marsocci, Valerio
Kopp, Niklas
Ramachandran, Rahul
Fraccaro, Paolo
Brunschwiler, Thomas
Cavallaro, Gabriele
Bernabe-Moreno, Juan
Longépé, Nicolas
contents We present TerraMind, the first any-to-any generative, multimodal foundation model for Earth observation (EO). Unlike other multimodal models, TerraMind is pretrained on dual-scale representations combining both token-level and pixel-level data across modalities. On a token level, TerraMind encodes high-level contextual information to learn cross-modal relationships, while on a pixel level, TerraMind leverages fine-grained representations to capture critical spatial nuances. We pretrained TerraMind on nine geospatial modalities of a global, large-scale dataset. In this paper, we demonstrate that (i) TerraMind's dual-scale early fusion approach unlocks a range of zero-shot and few-shot applications for Earth observation, (ii) TerraMind introduces "Thinking-in-Modalities" (TiM) -- the capability of generating additional artificial data during finetuning and inference to improve the model output -- and (iii) TerraMind achieves beyond state-of-the-art performance in community-standard benchmarks for EO like PANGAEA. The pretraining dataset, the model weights, and our code are open-sourced under a permissive license.
format Preprint
id arxiv_https___arxiv_org_abs_2504_11171
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle TerraMind: Large-Scale Generative Multimodality for Earth Observation
Jakubik, Johannes
Yang, Felix
Blumenstiel, Benedikt
Scheurer, Erik
Sedona, Rocco
Maurogiovanni, Stefano
Bosmans, Jente
Dionelis, Nikolaos
Marsocci, Valerio
Kopp, Niklas
Ramachandran, Rahul
Fraccaro, Paolo
Brunschwiler, Thomas
Cavallaro, Gabriele
Bernabe-Moreno, Juan
Longépé, Nicolas
Computer Vision and Pattern Recognition
Artificial Intelligence
We present TerraMind, the first any-to-any generative, multimodal foundation model for Earth observation (EO). Unlike other multimodal models, TerraMind is pretrained on dual-scale representations combining both token-level and pixel-level data across modalities. On a token level, TerraMind encodes high-level contextual information to learn cross-modal relationships, while on a pixel level, TerraMind leverages fine-grained representations to capture critical spatial nuances. We pretrained TerraMind on nine geospatial modalities of a global, large-scale dataset. In this paper, we demonstrate that (i) TerraMind's dual-scale early fusion approach unlocks a range of zero-shot and few-shot applications for Earth observation, (ii) TerraMind introduces "Thinking-in-Modalities" (TiM) -- the capability of generating additional artificial data during finetuning and inference to improve the model output -- and (iii) TerraMind achieves beyond state-of-the-art performance in community-standard benchmarks for EO like PANGAEA. The pretraining dataset, the model weights, and our code are open-sourced under a permissive license.
title TerraMind: Large-Scale Generative Multimodality for Earth Observation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2504.11171