One-step Diffusion with Distribution Matching Distillation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yin, Tianwei, Gharbi, Michaël, Zhang, Richard, Shechtman, Eli, Durand, Fredo, Freeman, William T., Park, Taesung
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908073886482432
author Yin, Tianwei
Gharbi, Michaël
Zhang, Richard
Shechtman, Eli
Durand, Fredo
Freeman, William T.
Park, Taesung
author_facet Yin, Tianwei
Gharbi, Michaël
Zhang, Richard
Shechtman, Eli
Durand, Fredo
Freeman, William T.
Park, Taesung
contents Diffusion models generate high-quality images but require dozens of forward passes. We introduce Distribution Matching Distillation (DMD), a procedure to transform a diffusion model into a one-step image generator with minimal impact on image quality. We enforce the one-step image generator match the diffusion model at distribution level, by minimizing an approximate KL divergence whose gradient can be expressed as the difference between 2 score functions, one of the target distribution and the other of the synthetic distribution being produced by our one-step generator. The score functions are parameterized as two diffusion models trained separately on each distribution. Combined with a simple regression loss matching the large-scale structure of the multi-step diffusion outputs, our method outperforms all published few-step diffusion approaches, reaching 2.62 FID on ImageNet 64x64 and 11.49 FID on zero-shot COCO-30k, comparable to Stable Diffusion but orders of magnitude faster. Utilizing FP16 inference, our model generates images at 20 FPS on modern hardware.
format Preprint
id arxiv_https___arxiv_org_abs_2311_18828
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle One-step Diffusion with Distribution Matching Distillation
Yin, Tianwei
Gharbi, Michaël
Zhang, Richard
Shechtman, Eli
Durand, Fredo
Freeman, William T.
Park, Taesung
Computer Vision and Pattern Recognition
Diffusion models generate high-quality images but require dozens of forward passes. We introduce Distribution Matching Distillation (DMD), a procedure to transform a diffusion model into a one-step image generator with minimal impact on image quality. We enforce the one-step image generator match the diffusion model at distribution level, by minimizing an approximate KL divergence whose gradient can be expressed as the difference between 2 score functions, one of the target distribution and the other of the synthetic distribution being produced by our one-step generator. The score functions are parameterized as two diffusion models trained separately on each distribution. Combined with a simple regression loss matching the large-scale structure of the multi-step diffusion outputs, our method outperforms all published few-step diffusion approaches, reaching 2.62 FID on ImageNet 64x64 and 11.49 FID on zero-shot COCO-30k, comparable to Stable Diffusion but orders of magnitude faster. Utilizing FP16 inference, our model generates images at 20 FPS on modern hardware.
title One-step Diffusion with Distribution Matching Distillation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2311.18828