Atomizer: Generalizing to new modalities by breaking satellite images down to a set of scalars

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: de Turckheim, Hugo Riffaud, Lobry, Sylvain, Interdonato, Roberto, Marcos, Diego
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915485115744256
author de Turckheim, Hugo Riffaud
Lobry, Sylvain
Interdonato, Roberto
Marcos, Diego
author_facet de Turckheim, Hugo Riffaud
Lobry, Sylvain
Interdonato, Roberto
Marcos, Diego
contents The growing number of Earth observation satellites has led to increasingly diverse remote sensing data, with varying spatial, spectral, and temporal configurations. Most existing models rely on fixed input formats and modality-specific encoders, which require retraining when new configurations are introduced, limiting their ability to generalize across modalities. We introduce Atomizer, a flexible architecture that represents remote sensing images as sets of scalars, each corresponding to a spectral band value of a pixel. Each scalar is enriched with contextual metadata (acquisition time, spatial resolution, wavelength, and bandwidth), producing an atomic representation that allows a single encoder to process arbitrary modalities without interpolation or resampling. Atomizer uses structured tokenization with Fourier features and non-uniform radial basis functions to encode content and context, and maps tokens into a latent space via cross-attention. Under modality-disjoint evaluations, Atomizer outperforms standard models and demonstrates robust performance across varying resolutions and spatial sizes.
format Preprint
id arxiv_https___arxiv_org_abs_2506_13542
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Atomizer: Generalizing to new modalities by breaking satellite images down to a set of scalars
de Turckheim, Hugo Riffaud
Lobry, Sylvain
Interdonato, Roberto
Marcos, Diego
Computer Vision and Pattern Recognition
The growing number of Earth observation satellites has led to increasingly diverse remote sensing data, with varying spatial, spectral, and temporal configurations. Most existing models rely on fixed input formats and modality-specific encoders, which require retraining when new configurations are introduced, limiting their ability to generalize across modalities. We introduce Atomizer, a flexible architecture that represents remote sensing images as sets of scalars, each corresponding to a spectral band value of a pixel. Each scalar is enriched with contextual metadata (acquisition time, spatial resolution, wavelength, and bandwidth), producing an atomic representation that allows a single encoder to process arbitrary modalities without interpolation or resampling. Atomizer uses structured tokenization with Fourier features and non-uniform radial basis functions to encode content and context, and maps tokens into a latent space via cross-attention. Under modality-disjoint evaluations, Atomizer outperforms standard models and demonstrates robust performance across varying resolutions and spatial sizes.
title Atomizer: Generalizing to new modalities by breaking satellite images down to a set of scalars
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.13542