Multi-modal contrastive learning adapts to intrinsic dimensions of shared latent variables

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gui, Yu, Ma, Cong, Ma, Zongming
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910951635156992
author Gui, Yu
Ma, Cong
Ma, Zongming
author_facet Gui, Yu
Ma, Cong
Ma, Zongming
contents Multi-modal contrastive learning as a self-supervised representation learning technique has achieved great success in foundation model training, such as CLIP~\citep{radford2021learning}. In this paper, we study the theoretical properties of the learned representations from multi-modal contrastive learning beyond linear representations and specific data distributions. Our analysis reveals that, enabled by temperature optimization, multi-modal contrastive learning not only maximizes mutual information between modalities but also adapts to intrinsic dimensions of data, which can be much lower than user-specified dimensions for representation vectors. Experiments on both synthetic and real-world datasets demonstrate the ability of contrastive learning to learn low-dimensional and informative representations, bridging theoretical insights and practical performance.
format Preprint
id arxiv_https___arxiv_org_abs_2505_12473
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Multi-modal contrastive learning adapts to intrinsic dimensions of shared latent variables
Gui, Yu
Ma, Cong
Ma, Zongming
Machine Learning
Statistics Theory
Multi-modal contrastive learning as a self-supervised representation learning technique has achieved great success in foundation model training, such as CLIP~\citep{radford2021learning}. In this paper, we study the theoretical properties of the learned representations from multi-modal contrastive learning beyond linear representations and specific data distributions. Our analysis reveals that, enabled by temperature optimization, multi-modal contrastive learning not only maximizes mutual information between modalities but also adapts to intrinsic dimensions of data, which can be much lower than user-specified dimensions for representation vectors. Experiments on both synthetic and real-world datasets demonstrate the ability of contrastive learning to learn low-dimensional and informative representations, bridging theoretical insights and practical performance.
title Multi-modal contrastive learning adapts to intrinsic dimensions of shared latent variables
topic Machine Learning
Statistics Theory
url https://arxiv.org/abs/2505.12473