DDIL: Diversity Enhancing Diffusion Distillation With Imitation Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Garrepalli, Risheek, Mahajan, Shweta, Hayat, Munawar, Porikli, Fatih
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915218200723456
author Garrepalli, Risheek
Mahajan, Shweta
Hayat, Munawar
Porikli, Fatih
author_facet Garrepalli, Risheek
Mahajan, Shweta
Hayat, Munawar
Porikli, Fatih
contents Diffusion models excel at generative modeling (e.g., text-to-image) but sampling requires multiple denoising network passes, limiting practicality. Efforts such as progressive distillation or consistency distillation have shown promise by reducing the number of passes at the expense of quality of the generated samples. In this work we identify co-variate shift as one of reason for poor performance of multi-step distilled models from compounding error at inference time. To address co-variate shift, we formulate diffusion distillation within imitation learning (DDIL) framework and enhance training distribution for distilling diffusion models on both data distribution (forward diffusion) and student induced distributions (backward diffusion). Training on data distribution helps to diversify the generations by preserving marginal data distribution and training on student distribution addresses compounding error by correcting covariate shift. In addition, we adopt reflected diffusion formulation for distillation and demonstrate improved performance, stable training across different distillation methods. We show that DDIL consistency improves on baseline algorithms of progressive distillation (PD), Latent consistency models (LCM) and Distribution Matching Distillation (DMD2).
format Preprint
id arxiv_https___arxiv_org_abs_2410_11971
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle DDIL: Diversity Enhancing Diffusion Distillation With Imitation Learning
Garrepalli, Risheek
Mahajan, Shweta
Hayat, Munawar
Porikli, Fatih
Machine Learning
Artificial Intelligence
Computer Vision and Pattern Recognition
Diffusion models excel at generative modeling (e.g., text-to-image) but sampling requires multiple denoising network passes, limiting practicality. Efforts such as progressive distillation or consistency distillation have shown promise by reducing the number of passes at the expense of quality of the generated samples. In this work we identify co-variate shift as one of reason for poor performance of multi-step distilled models from compounding error at inference time. To address co-variate shift, we formulate diffusion distillation within imitation learning (DDIL) framework and enhance training distribution for distilling diffusion models on both data distribution (forward diffusion) and student induced distributions (backward diffusion). Training on data distribution helps to diversify the generations by preserving marginal data distribution and training on student distribution addresses compounding error by correcting covariate shift. In addition, we adopt reflected diffusion formulation for distillation and demonstrate improved performance, stable training across different distillation methods. We show that DDIL consistency improves on baseline algorithms of progressive distillation (PD), Latent consistency models (LCM) and Distribution Matching Distillation (DMD2).
title DDIL: Diversity Enhancing Diffusion Distillation With Imitation Learning
topic Machine Learning
Artificial Intelligence
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2410.11971