Text-to-image Diffusion Models in Generative AI: A Survey

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Chenshuang, Zhang, Chaoning, Zhang, Mengchun, Kweon, In So, Kim, Junmo
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915008948994048
author Zhang, Chenshuang
Zhang, Chaoning
Zhang, Mengchun
Kweon, In So
Kim, Junmo
author_facet Zhang, Chenshuang
Zhang, Chaoning
Zhang, Mengchun
Kweon, In So
Kim, Junmo
contents This survey reviews the progress of diffusion models in generating images from text, ~\textit{i.e.} text-to-image diffusion models. As a self-contained work, this survey starts with a brief introduction of how diffusion models work for image synthesis, followed by the background for text-conditioned image synthesis. Based on that, we present an organized review of pioneering methods and their improvements on text-to-image generation. We further summarize applications beyond image generation, such as text-guided generation for various modalities like videos, and text-guided image editing. Beyond the progress made so far, we discuss existing challenges and promising future directions.
format Preprint
id arxiv_https___arxiv_org_abs_2303_07909
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Text-to-image Diffusion Models in Generative AI: A Survey
Zhang, Chenshuang
Zhang, Chaoning
Zhang, Mengchun
Kweon, In So
Kim, Junmo
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
This survey reviews the progress of diffusion models in generating images from text, ~\textit{i.e.} text-to-image diffusion models. As a self-contained work, this survey starts with a brief introduction of how diffusion models work for image synthesis, followed by the background for text-conditioned image synthesis. Based on that, we present an organized review of pioneering methods and their improvements on text-to-image generation. We further summarize applications beyond image generation, such as text-guided generation for various modalities like videos, and text-guided image editing. Beyond the progress made so far, we discuss existing challenges and promising future directions.
title Text-to-image Diffusion Models in Generative AI: A Survey
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2303.07909