Towards Controllable Image Generation through Representation-Conditioned Diffusion Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866911721396895744 |
|---|---|
| author | Karthikeyan, Nithesh Chandher Unger, Jonas Eilertsen, Gabriel |
| author_facet | Karthikeyan, Nithesh Chandher Unger, Jonas Eilertsen, Gabriel |
| contents | Diffusion models have emerged as powerful tools for high-quality image generation and editing, but guiding these models to produce specific outputs remains a challenge. Conventional approaches rely on conditioning mechanisms, such as text prompts or semantic maps, which require extensively annotated datasets. In this preliminary work, we explore diffusion models conditioned on representations from a pre-trained self-supervised model. The self-conditioning mechanism not only improves the quality of unconditional image generation, but also provides a representation space that can be used to control the generation. We explore this conditioning space by identifying directions of variations, and demonstrate promising properties in terms of smoothness and disentanglement. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2605_27343 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Towards Controllable Image Generation through Representation-Conditioned Diffusion Models Karthikeyan, Nithesh Chandher Unger, Jonas Eilertsen, Gabriel Computer Vision and Pattern Recognition Machine Learning Diffusion models have emerged as powerful tools for high-quality image generation and editing, but guiding these models to produce specific outputs remains a challenge. Conventional approaches rely on conditioning mechanisms, such as text prompts or semantic maps, which require extensively annotated datasets. In this preliminary work, we explore diffusion models conditioned on representations from a pre-trained self-supervised model. The self-conditioning mechanism not only improves the quality of unconditional image generation, but also provides a representation space that can be used to control the generation. We explore this conditioning space by identifying directions of variations, and demonstrate promising properties in terms of smoothness and disentanglement. |
| title | Towards Controllable Image Generation through Representation-Conditioned Diffusion Models |
| topic | Computer Vision and Pattern Recognition Machine Learning |
| url | https://arxiv.org/abs/2605.27343 |