A Noise is Worth Diffusion Guidance

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Ahn, Donghoon, Kang, Jiwon, Lee, Sanghyun, Min, Jaewon, Kim, Minjae, Jang, Wooseok, Cho, Hyoungwon, Paul, Sayak, Kim, SeonHwa, Cha, Eunju, Jin, Kyong Hwan, Kim, Seungryong
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912145091854336
author Ahn, Donghoon
Kang, Jiwon
Lee, Sanghyun
Min, Jaewon
Kim, Minjae
Jang, Wooseok
Cho, Hyoungwon
Paul, Sayak
Kim, SeonHwa
Cha, Eunju
Jin, Kyong Hwan
Kim, Seungryong
author_facet Ahn, Donghoon
Kang, Jiwon
Lee, Sanghyun
Min, Jaewon
Kim, Minjae
Jang, Wooseok
Cho, Hyoungwon
Paul, Sayak
Kim, SeonHwa
Cha, Eunju
Jin, Kyong Hwan
Kim, Seungryong
contents Diffusion models excel in generating high-quality images. However, current diffusion models struggle to produce reliable images without guidance methods, such as classifier-free guidance (CFG). Are guidance methods truly necessary? Observing that noise obtained via diffusion inversion can reconstruct high-quality images without guidance, we focus on the initial noise of the denoising pipeline. By mapping Gaussian noise to `guidance-free noise', we uncover that small low-magnitude low-frequency components significantly enhance the denoising process, removing the need for guidance and thus improving both inference throughput and memory. Expanding on this, we propose \ours, a novel method that replaces guidance methods with a single refinement of the initial noise. This refined noise enables high-quality image generation without guidance, within the same diffusion pipeline. Our noise-refining model leverages efficient noise-space learning, achieving rapid convergence and strong performance with just 50K text-image pairs. We validate its effectiveness across diverse metrics and analyze how refined noise can eliminate the need for guidance. See our project page: https://cvlab-kaist.github.io/NoiseRefine/.
format Preprint
id arxiv_https___arxiv_org_abs_2412_03895
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle A Noise is Worth Diffusion Guidance
Ahn, Donghoon
Kang, Jiwon
Lee, Sanghyun
Min, Jaewon
Kim, Minjae
Jang, Wooseok
Cho, Hyoungwon
Paul, Sayak
Kim, SeonHwa
Cha, Eunju
Jin, Kyong Hwan
Kim, Seungryong
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Diffusion models excel in generating high-quality images. However, current diffusion models struggle to produce reliable images without guidance methods, such as classifier-free guidance (CFG). Are guidance methods truly necessary? Observing that noise obtained via diffusion inversion can reconstruct high-quality images without guidance, we focus on the initial noise of the denoising pipeline. By mapping Gaussian noise to `guidance-free noise', we uncover that small low-magnitude low-frequency components significantly enhance the denoising process, removing the need for guidance and thus improving both inference throughput and memory. Expanding on this, we propose \ours, a novel method that replaces guidance methods with a single refinement of the initial noise. This refined noise enables high-quality image generation without guidance, within the same diffusion pipeline. Our noise-refining model leverages efficient noise-space learning, achieving rapid convergence and strong performance with just 50K text-image pairs. We validate its effectiveness across diverse metrics and analyze how refined noise can eliminate the need for guidance. See our project page: https://cvlab-kaist.github.io/NoiseRefine/.
title A Noise is Worth Diffusion Guidance
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2412.03895