Whitened CLIP as a Likelihood Surrogate of Images and Captions

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Betser, Roy, Levi, Meir Yossef, Gilboa, Guy
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918017499136000
author Betser, Roy
Levi, Meir Yossef
Gilboa, Guy
author_facet Betser, Roy
Levi, Meir Yossef
Gilboa, Guy
contents Likelihood approximations for images are not trivial to compute and can be useful in many applications. We examine the use of Contrastive Language-Image Pre-training (CLIP) to assess the likelihood of images and captions. We introduce \textit{Whitened CLIP}, a novel transformation of the CLIP latent space via an invertible linear operation. This transformation ensures that each feature in the embedding space has zero mean, unit standard deviation, and no correlation with all other features, resulting in an identity covariance matrix. We show that the whitened embeddings statistics can be well approximated as a standard normal distribution, thus, the log-likelihood is estimated simply by the square Euclidean norm in the whitened embedding space. The whitening procedure is completely training-free and performed using a pre-computed whitening matrix, hence, is very fast. We present several preliminary experiments demonstrating the properties and applicability of these likelihood scores to images and captions.
format Preprint
id arxiv_https___arxiv_org_abs_2505_06934
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Whitened CLIP as a Likelihood Surrogate of Images and Captions
Betser, Roy
Levi, Meir Yossef
Gilboa, Guy
Image and Video Processing
Computer Vision and Pattern Recognition
Likelihood approximations for images are not trivial to compute and can be useful in many applications. We examine the use of Contrastive Language-Image Pre-training (CLIP) to assess the likelihood of images and captions. We introduce \textit{Whitened CLIP}, a novel transformation of the CLIP latent space via an invertible linear operation. This transformation ensures that each feature in the embedding space has zero mean, unit standard deviation, and no correlation with all other features, resulting in an identity covariance matrix. We show that the whitened embeddings statistics can be well approximated as a standard normal distribution, thus, the log-likelihood is estimated simply by the square Euclidean norm in the whitened embedding space. The whitening procedure is completely training-free and performed using a pre-computed whitening matrix, hence, is very fast. We present several preliminary experiments demonstrating the properties and applicability of these likelihood scores to images and captions.
title Whitened CLIP as a Likelihood Surrogate of Images and Captions
topic Image and Video Processing
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.06934