Conditional Image Synthesis with Diffusion Models: A Survey

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhan, Zheyuan, Chen, Defang, Mei, Jian-Ping, Zhao, Zhenghe, Chen, Jiawei, Chen, Chun, Lyu, Siwei, Wang, Can
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912406996779008
author Zhan, Zheyuan
Chen, Defang
Mei, Jian-Ping
Zhao, Zhenghe
Chen, Jiawei
Chen, Chun
Lyu, Siwei
Wang, Can
author_facet Zhan, Zheyuan
Chen, Defang
Mei, Jian-Ping
Zhao, Zhenghe
Chen, Jiawei
Chen, Chun
Lyu, Siwei
Wang, Can
contents Conditional image synthesis based on user-specified requirements is a key component in creating complex visual content. In recent years, diffusion-based generative modeling has become a highly effective way for conditional image synthesis, leading to exponential growth in the literature. However, the complexity of diffusion-based modeling, the wide range of image synthesis tasks, and the diversity of conditioning mechanisms present significant challenges for researchers to keep up with rapid developments and to understand the core concepts on this topic. In this survey, we categorize existing works based on how conditions are integrated into the two fundamental components of diffusion-based modeling, $\textit{i.e.}$, the denoising network and the sampling process. We specifically highlight the underlying principles, advantages, and potential challenges of various conditioning approaches during the training, re-purposing, and specialization stages to construct a desired denoising network. We also summarize six mainstream conditioning mechanisms in the sampling process. All discussions are centered around popular applications. Finally, we pinpoint several critical yet still unsolved problems and suggest some possible solutions for future research. Our reviewed works are itemized at https://github.com/zju-pi/Awesome-Conditional-Diffusion-Models.
format Preprint
id arxiv_https___arxiv_org_abs_2409_19365
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Conditional Image Synthesis with Diffusion Models: A Survey
Zhan, Zheyuan
Chen, Defang
Mei, Jian-Ping
Zhao, Zhenghe
Chen, Jiawei
Chen, Chun
Lyu, Siwei
Wang, Can
Computer Vision and Pattern Recognition
Artificial Intelligence
Conditional image synthesis based on user-specified requirements is a key component in creating complex visual content. In recent years, diffusion-based generative modeling has become a highly effective way for conditional image synthesis, leading to exponential growth in the literature. However, the complexity of diffusion-based modeling, the wide range of image synthesis tasks, and the diversity of conditioning mechanisms present significant challenges for researchers to keep up with rapid developments and to understand the core concepts on this topic. In this survey, we categorize existing works based on how conditions are integrated into the two fundamental components of diffusion-based modeling, $\textit{i.e.}$, the denoising network and the sampling process. We specifically highlight the underlying principles, advantages, and potential challenges of various conditioning approaches during the training, re-purposing, and specialization stages to construct a desired denoising network. We also summarize six mainstream conditioning mechanisms in the sampling process. All discussions are centered around popular applications. Finally, we pinpoint several critical yet still unsolved problems and suggest some possible solutions for future research. Our reviewed works are itemized at https://github.com/zju-pi/Awesome-Conditional-Diffusion-Models.
title Conditional Image Synthesis with Diffusion Models: A Survey
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2409.19365