Can multimodal representation learning by alignment preserve modality-specific information?

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Thoreau, Romain, Levillain, Jessie, Derksen, Dawa
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866914051130392576
author Thoreau, Romain
Levillain, Jessie
Derksen, Dawa
author_facet Thoreau, Romain
Levillain, Jessie
Derksen, Dawa
contents Combining multimodal data is a key issue in a wide range of machine learning tasks, including many remote sensing problems. In Earth observation, early multimodal data fusion methods were based on specific neural network architectures and supervised learning. Ever since, the scarcity of labeled data has motivated self-supervised learning techniques. State-of-the-art multimodal representation learning techniques leverage the spatial alignment between satellite data from different modalities acquired over the same geographic area in order to foster a semantic alignment in the latent space. In this paper, we investigate how this methods can preserve task-relevant information that is not shared across modalities. First, we show, under simplifying assumptions, when alignment strategies fundamentally lead to an information loss. Then, we support our theoretical insight through numerical experiments in more realistic settings. With those theoretical and empirical evidences, we hope to support new developments in contrastive learning for the combination of multimodal satellite data. Our code and data is publicly available at https://github.com/Romain3Ch216/alg_maclean_25.
format Preprint
id arxiv_https___arxiv_org_abs_2509_17943
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Can multimodal representation learning by alignment preserve modality-specific information?
Thoreau, Romain
Levillain, Jessie
Derksen, Dawa
Computer Vision and Pattern Recognition
Machine Learning
Combining multimodal data is a key issue in a wide range of machine learning tasks, including many remote sensing problems. In Earth observation, early multimodal data fusion methods were based on specific neural network architectures and supervised learning. Ever since, the scarcity of labeled data has motivated self-supervised learning techniques. State-of-the-art multimodal representation learning techniques leverage the spatial alignment between satellite data from different modalities acquired over the same geographic area in order to foster a semantic alignment in the latent space. In this paper, we investigate how this methods can preserve task-relevant information that is not shared across modalities. First, we show, under simplifying assumptions, when alignment strategies fundamentally lead to an information loss. Then, we support our theoretical insight through numerical experiments in more realistic settings. With those theoretical and empirical evidences, we hope to support new developments in contrastive learning for the combination of multimodal satellite data. Our code and data is publicly available at https://github.com/Romain3Ch216/alg_maclean_25.
title Can multimodal representation learning by alignment preserve modality-specific information?
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2509.17943