Towards Controllable Image Generation through Representation-Conditioned Diffusion Models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Karthikeyan, Nithesh Chandher, Unger, Jonas, Eilertsen, Gabriel
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911721396895744
author Karthikeyan, Nithesh Chandher
Unger, Jonas
Eilertsen, Gabriel
author_facet Karthikeyan, Nithesh Chandher
Unger, Jonas
Eilertsen, Gabriel
contents Diffusion models have emerged as powerful tools for high-quality image generation and editing, but guiding these models to produce specific outputs remains a challenge. Conventional approaches rely on conditioning mechanisms, such as text prompts or semantic maps, which require extensively annotated datasets. In this preliminary work, we explore diffusion models conditioned on representations from a pre-trained self-supervised model. The self-conditioning mechanism not only improves the quality of unconditional image generation, but also provides a representation space that can be used to control the generation. We explore this conditioning space by identifying directions of variations, and demonstrate promising properties in terms of smoothness and disentanglement.
format Preprint
id arxiv_https___arxiv_org_abs_2605_27343
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Towards Controllable Image Generation through Representation-Conditioned Diffusion Models
Karthikeyan, Nithesh Chandher
Unger, Jonas
Eilertsen, Gabriel
Computer Vision and Pattern Recognition
Machine Learning
Diffusion models have emerged as powerful tools for high-quality image generation and editing, but guiding these models to produce specific outputs remains a challenge. Conventional approaches rely on conditioning mechanisms, such as text prompts or semantic maps, which require extensively annotated datasets. In this preliminary work, we explore diffusion models conditioned on representations from a pre-trained self-supervised model. The self-conditioning mechanism not only improves the quality of unconditional image generation, but also provides a representation space that can be used to control the generation. We explore this conditioning space by identifying directions of variations, and demonstrate promising properties in terms of smoothness and disentanglement.
title Towards Controllable Image Generation through Representation-Conditioned Diffusion Models
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2605.27343