One-step Diffusion with Distribution Matching Distillation
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866908073886482432 |
|---|---|
| author | Yin, Tianwei Gharbi, Michaël Zhang, Richard Shechtman, Eli Durand, Fredo Freeman, William T. Park, Taesung |
| author_facet | Yin, Tianwei Gharbi, Michaël Zhang, Richard Shechtman, Eli Durand, Fredo Freeman, William T. Park, Taesung |
| contents | Diffusion models generate high-quality images but require dozens of forward passes. We introduce Distribution Matching Distillation (DMD), a procedure to transform a diffusion model into a one-step image generator with minimal impact on image quality. We enforce the one-step image generator match the diffusion model at distribution level, by minimizing an approximate KL divergence whose gradient can be expressed as the difference between 2 score functions, one of the target distribution and the other of the synthetic distribution being produced by our one-step generator. The score functions are parameterized as two diffusion models trained separately on each distribution. Combined with a simple regression loss matching the large-scale structure of the multi-step diffusion outputs, our method outperforms all published few-step diffusion approaches, reaching 2.62 FID on ImageNet 64x64 and 11.49 FID on zero-shot COCO-30k, comparable to Stable Diffusion but orders of magnitude faster. Utilizing FP16 inference, our model generates images at 20 FPS on modern hardware. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2311_18828 |
| institution | arXiv |
| publishDate | 2023 |
| record_format | arxiv |
| spellingShingle | One-step Diffusion with Distribution Matching Distillation Yin, Tianwei Gharbi, Michaël Zhang, Richard Shechtman, Eli Durand, Fredo Freeman, William T. Park, Taesung Computer Vision and Pattern Recognition Diffusion models generate high-quality images but require dozens of forward passes. We introduce Distribution Matching Distillation (DMD), a procedure to transform a diffusion model into a one-step image generator with minimal impact on image quality. We enforce the one-step image generator match the diffusion model at distribution level, by minimizing an approximate KL divergence whose gradient can be expressed as the difference between 2 score functions, one of the target distribution and the other of the synthetic distribution being produced by our one-step generator. The score functions are parameterized as two diffusion models trained separately on each distribution. Combined with a simple regression loss matching the large-scale structure of the multi-step diffusion outputs, our method outperforms all published few-step diffusion approaches, reaching 2.62 FID on ImageNet 64x64 and 11.49 FID on zero-shot COCO-30k, comparable to Stable Diffusion but orders of magnitude faster. Utilizing FP16 inference, our model generates images at 20 FPS on modern hardware. |
| title | One-step Diffusion with Distribution Matching Distillation |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2311.18828 |