Learning to Discretize Denoising Diffusion ODEs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tong, Vinh, Trung-Dung, Hoang, Liu, Anji, Broeck, Guy Van den, Niepert, Mathias
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908370800214016
author Tong, Vinh
Trung-Dung, Hoang
Liu, Anji
Broeck, Guy Van den
Niepert, Mathias
author_facet Tong, Vinh
Trung-Dung, Hoang
Liu, Anji
Broeck, Guy Van den
Niepert, Mathias
contents Diffusion Probabilistic Models (DPMs) are generative models showing competitive performance in various domains, including image synthesis and 3D point cloud generation. Sampling from pre-trained DPMs involves multiple neural function evaluations (NFEs) to transform Gaussian noise samples into images, resulting in higher computational costs compared to single-step generative models such as GANs or VAEs. Therefore, reducing the number of NFEs while preserving generation quality is crucial. To address this, we propose LD3, a lightweight framework designed to learn the optimal time discretization for sampling. LD3 can be combined with various samplers and consistently improves generation quality without having to retrain resource-intensive neural networks. We demonstrate analytically and empirically that LD3 improves sampling efficiency with much less computational overhead. We evaluate our method with extensive experiments on 7 pre-trained models, covering unconditional and conditional sampling in both pixel-space and latent-space DPMs. We achieve FIDs of 2.38 (10 NFE), and 2.27 (10 NFE) on unconditional CIFAR10 and AFHQv2 in 5-10 minutes of training. LD3 offers an efficient approach to sampling from pre-trained diffusion models. Code is available at https://github.com/vinhsuhi/LD3.
format Preprint
id arxiv_https___arxiv_org_abs_2405_15506
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Learning to Discretize Denoising Diffusion ODEs
Tong, Vinh
Trung-Dung, Hoang
Liu, Anji
Broeck, Guy Van den
Niepert, Mathias
Machine Learning
Diffusion Probabilistic Models (DPMs) are generative models showing competitive performance in various domains, including image synthesis and 3D point cloud generation. Sampling from pre-trained DPMs involves multiple neural function evaluations (NFEs) to transform Gaussian noise samples into images, resulting in higher computational costs compared to single-step generative models such as GANs or VAEs. Therefore, reducing the number of NFEs while preserving generation quality is crucial. To address this, we propose LD3, a lightweight framework designed to learn the optimal time discretization for sampling. LD3 can be combined with various samplers and consistently improves generation quality without having to retrain resource-intensive neural networks. We demonstrate analytically and empirically that LD3 improves sampling efficiency with much less computational overhead. We evaluate our method with extensive experiments on 7 pre-trained models, covering unconditional and conditional sampling in both pixel-space and latent-space DPMs. We achieve FIDs of 2.38 (10 NFE), and 2.27 (10 NFE) on unconditional CIFAR10 and AFHQv2 in 5-10 minutes of training. LD3 offers an efficient approach to sampling from pre-trained diffusion models. Code is available at https://github.com/vinhsuhi/LD3.
title Learning to Discretize Denoising Diffusion ODEs
topic Machine Learning
url https://arxiv.org/abs/2405.15506