Flash Diffusion: Accelerating Any Conditional Diffusion Model for Few Steps Image Generation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Chadebec, Clément, Tasar, Onur, Benaroche, Eyal, Aubin, Benjamin
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912160506970112
author Chadebec, Clément
Tasar, Onur
Benaroche, Eyal
Aubin, Benjamin
author_facet Chadebec, Clément
Tasar, Onur
Benaroche, Eyal
Aubin, Benjamin
contents In this paper, we propose an efficient, fast, and versatile distillation method to accelerate the generation of pre-trained diffusion models: Flash Diffusion. The method reaches state-of-the-art performances in terms of FID and CLIP-Score for few steps image generation on the COCO2014 and COCO2017 datasets, while requiring only several GPU hours of training and fewer trainable parameters than existing methods. In addition to its efficiency, the versatility of the method is also exposed across several tasks such as text-to-image, inpainting, face-swapping, super-resolution and using different backbones such as UNet-based denoisers (SD1.5, SDXL) or DiT (Pixart-$α$), as well as adapters. In all cases, the method allowed to reduce drastically the number of sampling steps while maintaining very high-quality image generation. The official implementation is available at https://github.com/gojasper/flash-diffusion.
format Preprint
id arxiv_https___arxiv_org_abs_2406_02347
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Flash Diffusion: Accelerating Any Conditional Diffusion Model for Few Steps Image Generation
Chadebec, Clément
Tasar, Onur
Benaroche, Eyal
Aubin, Benjamin
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
In this paper, we propose an efficient, fast, and versatile distillation method to accelerate the generation of pre-trained diffusion models: Flash Diffusion. The method reaches state-of-the-art performances in terms of FID and CLIP-Score for few steps image generation on the COCO2014 and COCO2017 datasets, while requiring only several GPU hours of training and fewer trainable parameters than existing methods. In addition to its efficiency, the versatility of the method is also exposed across several tasks such as text-to-image, inpainting, face-swapping, super-resolution and using different backbones such as UNet-based denoisers (SD1.5, SDXL) or DiT (Pixart-$α$), as well as adapters. In all cases, the method allowed to reduce drastically the number of sampling steps while maintaining very high-quality image generation. The official implementation is available at https://github.com/gojasper/flash-diffusion.
title Flash Diffusion: Accelerating Any Conditional Diffusion Model for Few Steps Image Generation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2406.02347