Compress Guidance in Conditional Diffusion Sampling

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Dinh, Anh-Dung, Liu, Daochang, Xu, Chang
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929553609326592
author Dinh, Anh-Dung
Liu, Daochang
Xu, Chang
author_facet Dinh, Anh-Dung
Liu, Daochang
Xu, Chang
contents We found that enforcing guidance throughout the sampling process is often counterproductive due to the model-fitting issue, where samples are 'tuned' to match the classifier's parameters rather than generalizing the expected condition. This work identifies and quantifies the problem, demonstrating that reducing or excluding guidance at numerous timesteps can mitigate this issue. By distributing a small amount of guidance over a large number of sampling timesteps, we observe a significant improvement in image quality and diversity while also reducing the required guidance timesteps by nearly 40%. This approach addresses a major challenge in applying guidance effectively to generative tasks. Consequently, our proposed method, termed Compress Guidance, allows for the exclusion of a substantial number of guidance timesteps while still surpassing baseline models in image quality. We validate our approach through benchmarks on label-conditional and text-to-image generative tasks across various datasets and models.
format Preprint
id arxiv_https___arxiv_org_abs_2408_11194
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Compress Guidance in Conditional Diffusion Sampling
Dinh, Anh-Dung
Liu, Daochang
Xu, Chang
Computer Vision and Pattern Recognition
I.4
We found that enforcing guidance throughout the sampling process is often counterproductive due to the model-fitting issue, where samples are 'tuned' to match the classifier's parameters rather than generalizing the expected condition. This work identifies and quantifies the problem, demonstrating that reducing or excluding guidance at numerous timesteps can mitigate this issue. By distributing a small amount of guidance over a large number of sampling timesteps, we observe a significant improvement in image quality and diversity while also reducing the required guidance timesteps by nearly 40%. This approach addresses a major challenge in applying guidance effectively to generative tasks. Consequently, our proposed method, termed Compress Guidance, allows for the exclusion of a substantial number of guidance timesteps while still surpassing baseline models in image quality. We validate our approach through benchmarks on label-conditional and text-to-image generative tasks across various datasets and models.
title Compress Guidance in Conditional Diffusion Sampling
topic Computer Vision and Pattern Recognition
I.4
url https://arxiv.org/abs/2408.11194