Imagine Flash: Accelerating Emu Diffusion Models with Backward Distillation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Kohler, Jonas, Pumarola, Albert, Schönfeld, Edgar, Sanakoyeu, Artsiom, Sumbaly, Roshan, Vajda, Peter, Thabet, Ali
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909195719147520
author Kohler, Jonas
Pumarola, Albert
Schönfeld, Edgar
Sanakoyeu, Artsiom
Sumbaly, Roshan
Vajda, Peter
Thabet, Ali
author_facet Kohler, Jonas
Pumarola, Albert
Schönfeld, Edgar
Sanakoyeu, Artsiom
Sumbaly, Roshan
Vajda, Peter
Thabet, Ali
contents Diffusion models are a powerful generative framework, but come with expensive inference. Existing acceleration methods often compromise image quality or fail under complex conditioning when operating in an extremely low-step regime. In this work, we propose a novel distillation framework tailored to enable high-fidelity, diverse sample generation using just one to three steps. Our approach comprises three key components: (i) Backward Distillation, which mitigates training-inference discrepancies by calibrating the student on its own backward trajectory; (ii) Shifted Reconstruction Loss that dynamically adapts knowledge transfer based on the current time step; and (iii) Noise Correction, an inference-time technique that enhances sample quality by addressing singularities in noise prediction. Through extensive experiments, we demonstrate that our method outperforms existing competitors in quantitative metrics and human evaluations. Remarkably, it achieves performance comparable to the teacher model using only three denoising steps, enabling efficient high-quality generation.
format Preprint
id arxiv_https___arxiv_org_abs_2405_05224
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Imagine Flash: Accelerating Emu Diffusion Models with Backward Distillation
Kohler, Jonas
Pumarola, Albert
Schönfeld, Edgar
Sanakoyeu, Artsiom
Sumbaly, Roshan
Vajda, Peter
Thabet, Ali
Computer Vision and Pattern Recognition
Diffusion models are a powerful generative framework, but come with expensive inference. Existing acceleration methods often compromise image quality or fail under complex conditioning when operating in an extremely low-step regime. In this work, we propose a novel distillation framework tailored to enable high-fidelity, diverse sample generation using just one to three steps. Our approach comprises three key components: (i) Backward Distillation, which mitigates training-inference discrepancies by calibrating the student on its own backward trajectory; (ii) Shifted Reconstruction Loss that dynamically adapts knowledge transfer based on the current time step; and (iii) Noise Correction, an inference-time technique that enhances sample quality by addressing singularities in noise prediction. Through extensive experiments, we demonstrate that our method outperforms existing competitors in quantitative metrics and human evaluations. Remarkably, it achieves performance comparable to the teacher model using only three denoising steps, enabling efficient high-quality generation.
title Imagine Flash: Accelerating Emu Diffusion Models with Backward Distillation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2405.05224