Distilling Diffusion Models into Conditional GANs
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866917725286170624 |
|---|---|
| author | Kang, Minguk Zhang, Richard Barnes, Connelly Paris, Sylvain Kwak, Suha Park, Jaesik Shechtman, Eli Zhu, Jun-Yan Park, Taesung |
| author_facet | Kang, Minguk Zhang, Richard Barnes, Connelly Paris, Sylvain Kwak, Suha Park, Jaesik Shechtman, Eli Zhu, Jun-Yan Park, Taesung |
| contents | We propose a method to distill a complex multistep diffusion model into a single-step conditional GAN student model, dramatically accelerating inference, while preserving image quality. Our approach interprets diffusion distillation as a paired image-to-image translation task, using noise-to-image pairs of the diffusion model's ODE trajectory. For efficient regression loss computation, we propose E-LatentLPIPS, a perceptual loss operating directly in diffusion model's latent space, utilizing an ensemble of augmentations. Furthermore, we adapt a diffusion model to construct a multi-scale discriminator with a text alignment loss to build an effective conditional GAN-based formulation. E-LatentLPIPS converges more efficiently than many existing distillation methods, even accounting for dataset construction costs. We demonstrate that our one-step generator outperforms cutting-edge one-step diffusion distillation models -- DMD, SDXL-Turbo, and SDXL-Lightning -- on the zero-shot COCO benchmark. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2405_05967 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Distilling Diffusion Models into Conditional GANs Kang, Minguk Zhang, Richard Barnes, Connelly Paris, Sylvain Kwak, Suha Park, Jaesik Shechtman, Eli Zhu, Jun-Yan Park, Taesung Computer Vision and Pattern Recognition Graphics Machine Learning We propose a method to distill a complex multistep diffusion model into a single-step conditional GAN student model, dramatically accelerating inference, while preserving image quality. Our approach interprets diffusion distillation as a paired image-to-image translation task, using noise-to-image pairs of the diffusion model's ODE trajectory. For efficient regression loss computation, we propose E-LatentLPIPS, a perceptual loss operating directly in diffusion model's latent space, utilizing an ensemble of augmentations. Furthermore, we adapt a diffusion model to construct a multi-scale discriminator with a text alignment loss to build an effective conditional GAN-based formulation. E-LatentLPIPS converges more efficiently than many existing distillation methods, even accounting for dataset construction costs. We demonstrate that our one-step generator outperforms cutting-edge one-step diffusion distillation models -- DMD, SDXL-Turbo, and SDXL-Lightning -- on the zero-shot COCO benchmark. |
| title | Distilling Diffusion Models into Conditional GANs |
| topic | Computer Vision and Pattern Recognition Graphics Machine Learning |
| url | https://arxiv.org/abs/2405.05967 |