Quantizing Space and Time: Fusing Time Series and Images for Earth Observation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Basile, Gianfranco, Jakubik, Johannes, Blumenstiel, Benedikt, Brunschwiler, Thomas, Moreno, Juan Bernabe
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912844952371200
author Basile, Gianfranco
Jakubik, Johannes
Blumenstiel, Benedikt
Brunschwiler, Thomas
Moreno, Juan Bernabe
author_facet Basile, Gianfranco
Jakubik, Johannes
Blumenstiel, Benedikt
Brunschwiler, Thomas
Moreno, Juan Bernabe
contents We propose a task-agnostic framework for multimodal fusion of time series and single timestamp images, enabling cross-modal generation and robust downstream performance. Our approach explores deterministic and learned strategies for time series quantization and then leverages a masked correlation learning objective, aligning discrete image and time series tokens in a unified representation space. Instantiated in the Earth observation domain, the pretrained model generates consistent global temperature profiles from satellite imagery and is validated through counterfactual experiments. Across downstream tasks, our task-agnostic pretraining outperforms task-specific fusion by 6% in R^2 and 2% in RMSE on average, and exceeds baseline methods by 50% in R^2 and 12% in RMSE. Finally, we analyze gradient sensitivity across modalities, providing insights into model robustness. Code, data, and weights will be released under a permissive license.
format Preprint
id arxiv_https___arxiv_org_abs_2510_23118
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Quantizing Space and Time: Fusing Time Series and Images for Earth Observation
Basile, Gianfranco
Jakubik, Johannes
Blumenstiel, Benedikt
Brunschwiler, Thomas
Moreno, Juan Bernabe
Computer Vision and Pattern Recognition
We propose a task-agnostic framework for multimodal fusion of time series and single timestamp images, enabling cross-modal generation and robust downstream performance. Our approach explores deterministic and learned strategies for time series quantization and then leverages a masked correlation learning objective, aligning discrete image and time series tokens in a unified representation space. Instantiated in the Earth observation domain, the pretrained model generates consistent global temperature profiles from satellite imagery and is validated through counterfactual experiments. Across downstream tasks, our task-agnostic pretraining outperforms task-specific fusion by 6% in R^2 and 2% in RMSE on average, and exceeds baseline methods by 50% in R^2 and 12% in RMSE. Finally, we analyze gradient sensitivity across modalities, providing insights into model robustness. Code, data, and weights will be released under a permissive license.
title Quantizing Space and Time: Fusing Time Series and Images for Earth Observation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.23118