Decouple-Then-Merge: Finetune Diffusion Models as Multi-Task Learning

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Ma, Qianli, Ning, Xuefei, Liu, Dongrui, Niu, Li, Zhang, Linfeng
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909536982401024
author Ma, Qianli
Ning, Xuefei
Liu, Dongrui
Niu, Li
Zhang, Linfeng
author_facet Ma, Qianli
Ning, Xuefei
Liu, Dongrui
Niu, Li
Zhang, Linfeng
contents Diffusion models are trained by learning a sequence of models that reverse each step of noise corruption. Typically, the model parameters are fully shared across multiple timesteps to enhance training efficiency. However, since the denoising tasks differ at each timestep, the gradients computed at different timesteps may conflict, potentially degrading the overall performance of image generation. To solve this issue, this work proposes a \textbf{De}couple-then-\textbf{Me}rge (\textbf{DeMe}) framework, which begins with a pretrained model and finetunes separate models tailored to specific timesteps. We introduce several improved techniques during the finetuning stage to promote effective knowledge sharing while minimizing training interference across timesteps. Finally, after finetuning, these separate models can be merged into a single model in the parameter space, ensuring efficient and practical inference. Experimental results show significant generation quality improvements upon 6 benchmarks including Stable Diffusion on COCO30K, ImageNet1K, PartiPrompts, and DDPM on LSUN Church, LSUN Bedroom, and CIFAR10. Code is available at \href{https://github.com/MqLeet/DeMe}{GitHub}.
format Preprint
id arxiv_https___arxiv_org_abs_2410_06664
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Decouple-Then-Merge: Finetune Diffusion Models as Multi-Task Learning
Ma, Qianli
Ning, Xuefei
Liu, Dongrui
Niu, Li
Zhang, Linfeng
Computer Vision and Pattern Recognition
Artificial Intelligence
Diffusion models are trained by learning a sequence of models that reverse each step of noise corruption. Typically, the model parameters are fully shared across multiple timesteps to enhance training efficiency. However, since the denoising tasks differ at each timestep, the gradients computed at different timesteps may conflict, potentially degrading the overall performance of image generation. To solve this issue, this work proposes a \textbf{De}couple-then-\textbf{Me}rge (\textbf{DeMe}) framework, which begins with a pretrained model and finetunes separate models tailored to specific timesteps. We introduce several improved techniques during the finetuning stage to promote effective knowledge sharing while minimizing training interference across timesteps. Finally, after finetuning, these separate models can be merged into a single model in the parameter space, ensuring efficient and practical inference. Experimental results show significant generation quality improvements upon 6 benchmarks including Stable Diffusion on COCO30K, ImageNet1K, PartiPrompts, and DDPM on LSUN Church, LSUN Bedroom, and CIFAR10. Code is available at \href{https://github.com/MqLeet/DeMe}{GitHub}.
title Decouple-Then-Merge: Finetune Diffusion Models as Multi-Task Learning
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2410.06664