Distilling Diffusion Models into Conditional GANs

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Kang, Minguk, Zhang, Richard, Barnes, Connelly, Paris, Sylvain, Kwak, Suha, Park, Jaesik, Shechtman, Eli, Zhu, Jun-Yan, Park, Taesung
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866917725286170624
author Kang, Minguk
Zhang, Richard
Barnes, Connelly
Paris, Sylvain
Kwak, Suha
Park, Jaesik
Shechtman, Eli
Zhu, Jun-Yan
Park, Taesung
author_facet Kang, Minguk
Zhang, Richard
Barnes, Connelly
Paris, Sylvain
Kwak, Suha
Park, Jaesik
Shechtman, Eli
Zhu, Jun-Yan
Park, Taesung
contents We propose a method to distill a complex multistep diffusion model into a single-step conditional GAN student model, dramatically accelerating inference, while preserving image quality. Our approach interprets diffusion distillation as a paired image-to-image translation task, using noise-to-image pairs of the diffusion model's ODE trajectory. For efficient regression loss computation, we propose E-LatentLPIPS, a perceptual loss operating directly in diffusion model's latent space, utilizing an ensemble of augmentations. Furthermore, we adapt a diffusion model to construct a multi-scale discriminator with a text alignment loss to build an effective conditional GAN-based formulation. E-LatentLPIPS converges more efficiently than many existing distillation methods, even accounting for dataset construction costs. We demonstrate that our one-step generator outperforms cutting-edge one-step diffusion distillation models -- DMD, SDXL-Turbo, and SDXL-Lightning -- on the zero-shot COCO benchmark.
format Preprint
id arxiv_https___arxiv_org_abs_2405_05967
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Distilling Diffusion Models into Conditional GANs
Kang, Minguk
Zhang, Richard
Barnes, Connelly
Paris, Sylvain
Kwak, Suha
Park, Jaesik
Shechtman, Eli
Zhu, Jun-Yan
Park, Taesung
Computer Vision and Pattern Recognition
Graphics
Machine Learning
We propose a method to distill a complex multistep diffusion model into a single-step conditional GAN student model, dramatically accelerating inference, while preserving image quality. Our approach interprets diffusion distillation as a paired image-to-image translation task, using noise-to-image pairs of the diffusion model's ODE trajectory. For efficient regression loss computation, we propose E-LatentLPIPS, a perceptual loss operating directly in diffusion model's latent space, utilizing an ensemble of augmentations. Furthermore, we adapt a diffusion model to construct a multi-scale discriminator with a text alignment loss to build an effective conditional GAN-based formulation. E-LatentLPIPS converges more efficiently than many existing distillation methods, even accounting for dataset construction costs. We demonstrate that our one-step generator outperforms cutting-edge one-step diffusion distillation models -- DMD, SDXL-Turbo, and SDXL-Lightning -- on the zero-shot COCO benchmark.
title Distilling Diffusion Models into Conditional GANs
topic Computer Vision and Pattern Recognition
Graphics
Machine Learning
url https://arxiv.org/abs/2405.05967