Guiding a Diffusion Model with a Bad Version of Itself

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Karras, Tero, Aittala, Miika, Kynkäänniemi, Tuomas, Lehtinen, Jaakko, Aila, Timo, Laine, Samuli
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915070591631360
author Karras, Tero
Aittala, Miika
Kynkäänniemi, Tuomas
Lehtinen, Jaakko
Aila, Timo
Laine, Samuli
author_facet Karras, Tero
Aittala, Miika
Kynkäänniemi, Tuomas
Lehtinen, Jaakko
Aila, Timo
Laine, Samuli
contents The primary axes of interest in image-generating diffusion models are image quality, the amount of variation in the results, and how well the results align with a given condition, e.g., a class label or a text prompt. The popular classifier-free guidance approach uses an unconditional model to guide a conditional model, leading to simultaneously better prompt alignment and higher-quality images at the cost of reduced variation. These effects seem inherently entangled, and thus hard to control. We make the surprising observation that it is possible to obtain disentangled control over image quality without compromising the amount of variation by guiding generation using a smaller, less-trained version of the model itself rather than an unconditional model. This leads to significant improvements in ImageNet generation, setting record FIDs of 1.01 for 64x64 and 1.25 for 512x512, using publicly available networks. Furthermore, the method is also applicable to unconditional diffusion models, drastically improving their quality.
format Preprint
id arxiv_https___arxiv_org_abs_2406_02507
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Guiding a Diffusion Model with a Bad Version of Itself
Karras, Tero
Aittala, Miika
Kynkäänniemi, Tuomas
Lehtinen, Jaakko
Aila, Timo
Laine, Samuli
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Neural and Evolutionary Computing
The primary axes of interest in image-generating diffusion models are image quality, the amount of variation in the results, and how well the results align with a given condition, e.g., a class label or a text prompt. The popular classifier-free guidance approach uses an unconditional model to guide a conditional model, leading to simultaneously better prompt alignment and higher-quality images at the cost of reduced variation. These effects seem inherently entangled, and thus hard to control. We make the surprising observation that it is possible to obtain disentangled control over image quality without compromising the amount of variation by guiding generation using a smaller, less-trained version of the model itself rather than an unconditional model. This leads to significant improvements in ImageNet generation, setting record FIDs of 1.01 for 64x64 and 1.25 for 512x512, using publicly available networks. Furthermore, the method is also applicable to unconditional diffusion models, drastically improving their quality.
title Guiding a Diffusion Model with a Bad Version of Itself
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Neural and Evolutionary Computing
url https://arxiv.org/abs/2406.02507