Global Minimizers of Sigmoid Contrastive Loss

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bangachev, Kiril, Bresler, Guy, Noman, Iliyas, Polyanskiy, Yury
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912959720062976
author Bangachev, Kiril
Bresler, Guy
Noman, Iliyas
Polyanskiy, Yury
author_facet Bangachev, Kiril
Bresler, Guy
Noman, Iliyas
Polyanskiy, Yury
contents The meta-task of obtaining and aligning representations through contrastive pretraining is steadily gaining importance since its introduction in CLIP and ALIGN. In this paper we theoretically explain the advantages of synchronizing with trainable inverse temperature and bias under the sigmoid loss, as implemented in the recent SigLIP and SigLIP2 models of Google DeepMind. Temperature and bias can drive the loss function to zero for a rich class of configurations that we call $(\mathsf{m}, \mathsf{b}_{\mathsf{rel}})$-Constellations. $(\mathsf{m}, \mathsf{b}_{\mathsf{rel}})$-Constellations are a novel combinatorial object related to spherical codes and are parametrized by a margin $\mathsf{m}$ and relative bias $\mathsf{b}_{\mathsf{rel}}$. We use our characterization of constellations to theoretically justify the success of SigLIP on retrieval, to explain the modality gap present in SigLIP and CLIP, and to identify the necessary dimension for producing high-quality representations. Finally, we propose a reparameterization of the sigmoid loss with explicit relative bias, which improves training dynamics in experiments with synthetic data.
format Preprint
id arxiv_https___arxiv_org_abs_2509_18552
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Global Minimizers of Sigmoid Contrastive Loss
Bangachev, Kiril
Bresler, Guy
Noman, Iliyas
Polyanskiy, Yury
Machine Learning
Artificial Intelligence
The meta-task of obtaining and aligning representations through contrastive pretraining is steadily gaining importance since its introduction in CLIP and ALIGN. In this paper we theoretically explain the advantages of synchronizing with trainable inverse temperature and bias under the sigmoid loss, as implemented in the recent SigLIP and SigLIP2 models of Google DeepMind. Temperature and bias can drive the loss function to zero for a rich class of configurations that we call $(\mathsf{m}, \mathsf{b}_{\mathsf{rel}})$-Constellations. $(\mathsf{m}, \mathsf{b}_{\mathsf{rel}})$-Constellations are a novel combinatorial object related to spherical codes and are parametrized by a margin $\mathsf{m}$ and relative bias $\mathsf{b}_{\mathsf{rel}}$. We use our characterization of constellations to theoretically justify the success of SigLIP on retrieval, to explain the modality gap present in SigLIP and CLIP, and to identify the necessary dimension for producing high-quality representations. Finally, we propose a reparameterization of the sigmoid loss with explicit relative bias, which improves training dynamics in experiments with synthetic data.
title Global Minimizers of Sigmoid Contrastive Loss
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2509.18552