Diffusion Models Need Visual Priors for Image Generation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Yue, Xiaoyu, Wang, Zidong, Lu, Zeyu, Sun, Shuyang, Wei, Meng, Ouyang, Wanli, Bai, Lei, Zhou, Luping
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912069028151296
author Yue, Xiaoyu
Wang, Zidong
Lu, Zeyu
Sun, Shuyang
Wei, Meng
Ouyang, Wanli
Bai, Lei
Zhou, Luping
author_facet Yue, Xiaoyu
Wang, Zidong
Lu, Zeyu
Sun, Shuyang
Wei, Meng
Ouyang, Wanli
Bai, Lei
Zhou, Luping
contents Conventional class-guided diffusion models generally succeed in generating images with correct semantic content, but often struggle with texture details. This limitation stems from the usage of class priors, which only provide coarse and limited conditional information. To address this issue, we propose Diffusion on Diffusion (DoD), an innovative multi-stage generation framework that first extracts visual priors from previously generated samples, then provides rich guidance for the diffusion model leveraging visual priors from the early stages of diffusion sampling. Specifically, we introduce a latent embedding module that employs a compression-reconstruction approach to discard redundant detail information from the conditional samples in each stage, retaining only the semantic information for guidance. We evaluate DoD on the popular ImageNet-$256 \times 256$ dataset, reducing 7$\times$ training cost compared to SiT and DiT with even better performance in terms of the FID-50K score. Our largest model DoD-XL achieves an FID-50K score of 1.83 with only 1 million training steps, which surpasses other state-of-the-art methods without bells and whistles during inference.
format Preprint
id arxiv_https___arxiv_org_abs_2410_08531
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Diffusion Models Need Visual Priors for Image Generation
Yue, Xiaoyu
Wang, Zidong
Lu, Zeyu
Sun, Shuyang
Wei, Meng
Ouyang, Wanli
Bai, Lei
Zhou, Luping
Computer Vision and Pattern Recognition
Conventional class-guided diffusion models generally succeed in generating images with correct semantic content, but often struggle with texture details. This limitation stems from the usage of class priors, which only provide coarse and limited conditional information. To address this issue, we propose Diffusion on Diffusion (DoD), an innovative multi-stage generation framework that first extracts visual priors from previously generated samples, then provides rich guidance for the diffusion model leveraging visual priors from the early stages of diffusion sampling. Specifically, we introduce a latent embedding module that employs a compression-reconstruction approach to discard redundant detail information from the conditional samples in each stage, retaining only the semantic information for guidance. We evaluate DoD on the popular ImageNet-$256 \times 256$ dataset, reducing 7$\times$ training cost compared to SiT and DiT with even better performance in terms of the FID-50K score. Our largest model DoD-XL achieves an FID-50K score of 1.83 with only 1 million training steps, which surpasses other state-of-the-art methods without bells and whistles during inference.
title Diffusion Models Need Visual Priors for Image Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2410.08531