Efficient Diffusion Training via Min-SNR Weighting Strategy

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Hang, Tiankai, Gu, Shuyang, Li, Chen, Bao, Jianmin, Chen, Dong, Hu, Han, Geng, Xin, Guo, Baining
Format: Preprint
Veröffentlicht: 2023
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866910359618584576
author Hang, Tiankai
Gu, Shuyang
Li, Chen
Bao, Jianmin
Chen, Dong
Hu, Han
Geng, Xin
Guo, Baining
author_facet Hang, Tiankai
Gu, Shuyang
Li, Chen
Bao, Jianmin
Chen, Dong
Hu, Han
Geng, Xin
Guo, Baining
contents Denoising diffusion models have been a mainstream approach for image generation, however, training these models often suffers from slow convergence. In this paper, we discovered that the slow convergence is partly due to conflicting optimization directions between timesteps. To address this issue, we treat the diffusion training as a multi-task learning problem, and introduce a simple yet effective approach referred to as Min-SNR-$γ$. This method adapts loss weights of timesteps based on clamped signal-to-noise ratios, which effectively balances the conflicts among timesteps. Our results demonstrate a significant improvement in converging speed, 3.4$\times$ faster than previous weighting strategies. It is also more effective, achieving a new record FID score of 2.06 on the ImageNet $256\times256$ benchmark using smaller architectures than that employed in previous state-of-the-art. The code is available at https://github.com/TiankaiHang/Min-SNR-Diffusion-Training.
format Preprint
id arxiv_https___arxiv_org_abs_2303_09556
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Efficient Diffusion Training via Min-SNR Weighting Strategy
Hang, Tiankai
Gu, Shuyang
Li, Chen
Bao, Jianmin
Chen, Dong
Hu, Han
Geng, Xin
Guo, Baining
Computer Vision and Pattern Recognition
Denoising diffusion models have been a mainstream approach for image generation, however, training these models often suffers from slow convergence. In this paper, we discovered that the slow convergence is partly due to conflicting optimization directions between timesteps. To address this issue, we treat the diffusion training as a multi-task learning problem, and introduce a simple yet effective approach referred to as Min-SNR-$γ$. This method adapts loss weights of timesteps based on clamped signal-to-noise ratios, which effectively balances the conflicts among timesteps. Our results demonstrate a significant improvement in converging speed, 3.4$\times$ faster than previous weighting strategies. It is also more effective, achieving a new record FID score of 2.06 on the ImageNet $256\times256$ benchmark using smaller architectures than that employed in previous state-of-the-art. The code is available at https://github.com/TiankaiHang/Min-SNR-Diffusion-Training.
title Efficient Diffusion Training via Min-SNR Weighting Strategy
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2303.09556