Implicit Inversion turns CLIP into a Decoder

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: D'Orazio, Antonio, Briglia, Maria Rosaria, Crisostomi, Donato, Loi, Dario, Rodolà, Emanuele, Masi, Iacopo
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909636426203136
author D'Orazio, Antonio
Briglia, Maria Rosaria
Crisostomi, Donato
Loi, Dario
Rodolà, Emanuele
Masi, Iacopo
author_facet D'Orazio, Antonio
Briglia, Maria Rosaria
Crisostomi, Donato
Loi, Dario
Rodolà, Emanuele
Masi, Iacopo
contents CLIP is a discriminative model trained to align images and text in a shared embedding space. Due to its multimodal structure, it serves as the backbone of many generative pipelines, where a decoder is trained to map from the shared space back to images. In this work, we show that image synthesis is nevertheless possible using CLIP alone -- without any decoder, training, or fine-tuning. Our approach optimizes a frequency-aware implicit neural representation that encourages coarse-to-fine generation by stratifying frequencies across network layers. To stabilize this inverse mapping, we introduce adversarially robust initialization, a lightweight Orthogonal Procrustes projection to align local text and image embeddings, and a blending loss that anchors outputs to natural image statistics. Without altering CLIP's weights, this framework unlocks capabilities such as text-to-image generation, style transfer, and image reconstruction. These findings suggest that discriminative models may hold untapped generative potential, hidden in plain sight.
format Preprint
id arxiv_https___arxiv_org_abs_2505_23161
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Implicit Inversion turns CLIP into a Decoder
D'Orazio, Antonio
Briglia, Maria Rosaria
Crisostomi, Donato
Loi, Dario
Rodolà, Emanuele
Masi, Iacopo
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
CLIP is a discriminative model trained to align images and text in a shared embedding space. Due to its multimodal structure, it serves as the backbone of many generative pipelines, where a decoder is trained to map from the shared space back to images. In this work, we show that image synthesis is nevertheless possible using CLIP alone -- without any decoder, training, or fine-tuning. Our approach optimizes a frequency-aware implicit neural representation that encourages coarse-to-fine generation by stratifying frequencies across network layers. To stabilize this inverse mapping, we introduce adversarially robust initialization, a lightweight Orthogonal Procrustes projection to align local text and image embeddings, and a blending loss that anchors outputs to natural image statistics. Without altering CLIP's weights, this framework unlocks capabilities such as text-to-image generation, style transfer, and image reconstruction. These findings suggest that discriminative models may hold untapped generative potential, hidden in plain sight.
title Implicit Inversion turns CLIP into a Decoder
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2505.23161