A Closer Look at Multimodal Representation Collapse

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chaudhuri, Abhra, Dutta, Anjan, Bui, Tu, Georgescu, Serban
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913992006434816
author Chaudhuri, Abhra
Dutta, Anjan
Bui, Tu
Georgescu, Serban
author_facet Chaudhuri, Abhra
Dutta, Anjan
Bui, Tu
Georgescu, Serban
contents We aim to develop a fundamental understanding of modality collapse, a recently observed empirical phenomenon wherein models trained for multimodal fusion tend to rely only on a subset of the modalities, ignoring the rest. We show that modality collapse happens when noisy features from one modality are entangled, via a shared set of neurons in the fusion head, with predictive features from another, effectively masking out positive contributions from the predictive features of the former modality and leading to its collapse. We further prove that cross-modal knowledge distillation implicitly disentangles such representations by freeing up rank bottlenecks in the student encoder, denoising the fusion-head outputs without negatively impacting the predictive features from either modality. Based on the above findings, we propose an algorithm that prevents modality collapse through explicit basis reallocation, with applications in dealing with missing modalities. Extensive experiments on multiple multimodal benchmarks validate our theoretical claims. Project page: https://abhrac.github.io/mmcollapse/.
format Preprint
id arxiv_https___arxiv_org_abs_2505_22483
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Closer Look at Multimodal Representation Collapse
Chaudhuri, Abhra
Dutta, Anjan
Bui, Tu
Georgescu, Serban
Machine Learning
Artificial Intelligence
Computer Vision and Pattern Recognition
We aim to develop a fundamental understanding of modality collapse, a recently observed empirical phenomenon wherein models trained for multimodal fusion tend to rely only on a subset of the modalities, ignoring the rest. We show that modality collapse happens when noisy features from one modality are entangled, via a shared set of neurons in the fusion head, with predictive features from another, effectively masking out positive contributions from the predictive features of the former modality and leading to its collapse. We further prove that cross-modal knowledge distillation implicitly disentangles such representations by freeing up rank bottlenecks in the student encoder, denoising the fusion-head outputs without negatively impacting the predictive features from either modality. Based on the above findings, we propose an algorithm that prevents modality collapse through explicit basis reallocation, with applications in dealing with missing modalities. Extensive experiments on multiple multimodal benchmarks validate our theoretical claims. Project page: https://abhrac.github.io/mmcollapse/.
title A Closer Look at Multimodal Representation Collapse
topic Machine Learning
Artificial Intelligence
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.22483