Adapting Self-Supervised Representations as a Latent Space for Efficient Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gui, Ming, Schusterbauer, Johannes, Phan, Timy, Krause, Felix, Susskind, Josh, Bautista, Miguel Angel, Ommer, Björn
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908981499265024
author Gui, Ming
Schusterbauer, Johannes
Phan, Timy
Krause, Felix
Susskind, Josh
Bautista, Miguel Angel
Ommer, Björn
author_facet Gui, Ming
Schusterbauer, Johannes
Phan, Timy
Krause, Felix
Susskind, Josh
Bautista, Miguel Angel
Ommer, Björn
contents We introduce Representation Tokenizer (RepTok), a generative modeling framework that represents an image using a single continuous latent token obtained from self-supervised vision transformers. Building on a pre-trained SSL encoder, we fine-tune only the semantic token embedding and pair it with a generative decoder trained jointly using a standard flow matching objective. This adaptation enriches the token with low-level, reconstruction-relevant details, enabling faithful image reconstruction. To preserve the favorable geometry of the original SSL space, we add a cosine-similarity loss that regularizes the adapted token, ensuring the latent space remains smooth and suitable for generation. Our single-token formulation resolves spatial redundancies of 2D latent spaces and significantly reduces training costs. Despite its simplicity and efficiency, RepTok achieves competitive results on class-conditional ImageNet generation and naturally extends to text-to-image synthesis, reaching competitive zero-shot performance on MS-COCO under extremely limited training budgets. Our findings highlight the potential of fine-tuned SSL representations as compact and effective latent spaces for efficient generative modeling.
format Preprint
id arxiv_https___arxiv_org_abs_2510_14630
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Adapting Self-Supervised Representations as a Latent Space for Efficient Generation
Gui, Ming
Schusterbauer, Johannes
Phan, Timy
Krause, Felix
Susskind, Josh
Bautista, Miguel Angel
Ommer, Björn
Computer Vision and Pattern Recognition
We introduce Representation Tokenizer (RepTok), a generative modeling framework that represents an image using a single continuous latent token obtained from self-supervised vision transformers. Building on a pre-trained SSL encoder, we fine-tune only the semantic token embedding and pair it with a generative decoder trained jointly using a standard flow matching objective. This adaptation enriches the token with low-level, reconstruction-relevant details, enabling faithful image reconstruction. To preserve the favorable geometry of the original SSL space, we add a cosine-similarity loss that regularizes the adapted token, ensuring the latent space remains smooth and suitable for generation. Our single-token formulation resolves spatial redundancies of 2D latent spaces and significantly reduces training costs. Despite its simplicity and efficiency, RepTok achieves competitive results on class-conditional ImageNet generation and naturally extends to text-to-image synthesis, reaching competitive zero-shot performance on MS-COCO under extremely limited training budgets. Our findings highlight the potential of fine-tuned SSL representations as compact and effective latent spaces for efficient generative modeling.
title Adapting Self-Supervised Representations as a Latent Space for Efficient Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.14630