Towards More Accurate Diffusion Model Acceleration with A Timestep Tuner

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Xia, Mengfei, Shen, Yujun, Lei, Changsong, Zhou, Yu, Yi, Ran, Zhao, Deli, Wang, Wenping, Liu, Yong-Jin
Format: Preprint
Veröffentlicht: 2023
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908569823084544
author Xia, Mengfei
Shen, Yujun
Lei, Changsong
Zhou, Yu
Yi, Ran
Zhao, Deli
Wang, Wenping
Liu, Yong-Jin
author_facet Xia, Mengfei
Shen, Yujun
Lei, Changsong
Zhou, Yu
Yi, Ran
Zhao, Deli
Wang, Wenping
Liu, Yong-Jin
contents A diffusion model, which is formulated to produce an image using thousands of denoising steps, usually suffers from a slow inference speed. Existing acceleration algorithms simplify the sampling by skipping most steps yet exhibit considerable performance degradation. By viewing the generation of diffusion models as a discretized integral process, we argue that the quality drop is partly caused by applying an inaccurate integral direction to a timestep interval. To rectify this issue, we propose a \textbf{timestep tuner} that helps find a more accurate integral direction for a particular interval at the minimum cost. Specifically, at each denoising step, we replace the original parameterization by conditioning the network on a new timestep, enforcing the sampling distribution towards the real one. Extensive experiments show that our plug-in design can be trained efficiently and boost the inference performance of various state-of-the-art acceleration methods, especially when there are few denoising steps. For example, when using 10 denoising steps on LSUN Bedroom dataset, we improve the FID of DDIM from 9.65 to 6.07, simply by adopting our method for a more appropriate set of timesteps. Code is available at \href{https://github.com/THU-LYJ-Lab/time-tuner}{https://github.com/THU-LYJ-Lab/time-tuner}.
format Preprint
id arxiv_https___arxiv_org_abs_2310_09469
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Towards More Accurate Diffusion Model Acceleration with A Timestep Tuner
Xia, Mengfei
Shen, Yujun
Lei, Changsong
Zhou, Yu
Yi, Ran
Zhao, Deli
Wang, Wenping
Liu, Yong-Jin
Computer Vision and Pattern Recognition
A diffusion model, which is formulated to produce an image using thousands of denoising steps, usually suffers from a slow inference speed. Existing acceleration algorithms simplify the sampling by skipping most steps yet exhibit considerable performance degradation. By viewing the generation of diffusion models as a discretized integral process, we argue that the quality drop is partly caused by applying an inaccurate integral direction to a timestep interval. To rectify this issue, we propose a \textbf{timestep tuner} that helps find a more accurate integral direction for a particular interval at the minimum cost. Specifically, at each denoising step, we replace the original parameterization by conditioning the network on a new timestep, enforcing the sampling distribution towards the real one. Extensive experiments show that our plug-in design can be trained efficiently and boost the inference performance of various state-of-the-art acceleration methods, especially when there are few denoising steps. For example, when using 10 denoising steps on LSUN Bedroom dataset, we improve the FID of DDIM from 9.65 to 6.07, simply by adopting our method for a more appropriate set of timesteps. Code is available at \href{https://github.com/THU-LYJ-Lab/time-tuner}{https://github.com/THU-LYJ-Lab/time-tuner}.
title Towards More Accurate Diffusion Model Acceleration with A Timestep Tuner
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2310.09469