Text-to-Image Cross-Modal Generation: A Systematic Review

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Żelaszczyk, Maciej, Mańdziuk, Jacek
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911762164482048
author Żelaszczyk, Maciej
Mańdziuk, Jacek
author_facet Żelaszczyk, Maciej
Mańdziuk, Jacek
contents We review research on generating visual data from text from the angle of "cross-modal generation." This point of view allows us to draw parallels between various methods geared towards working on input text and producing visual output, without limiting the analysis to narrow sub-areas. It also results in the identification of common templates in the field, which are then compared and contrasted both within pools of similar methods and across lines of research. We provide a breakdown of text-to-image generation into various flavors of image-from-text methods, video-from-text methods, image editing, self-supervised and graph-based approaches. In this discussion, we focus on research papers published at 8 leading machine learning conferences in the years 2016-2022, also incorporating a number of relevant papers not matching the outlined search criteria. The conducted review suggests a significant increase in the number of papers published in the area and highlights research gaps and potential lines of investigation. To our knowledge, this is the first review to systematically look at text-to-image generation from the perspective of "cross-modal generation."
format Preprint
id arxiv_https___arxiv_org_abs_2401_11631
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Text-to-Image Cross-Modal Generation: A Systematic Review
Żelaszczyk, Maciej
Mańdziuk, Jacek
Computer Vision and Pattern Recognition
Computation and Language
Machine Learning
We review research on generating visual data from text from the angle of "cross-modal generation." This point of view allows us to draw parallels between various methods geared towards working on input text and producing visual output, without limiting the analysis to narrow sub-areas. It also results in the identification of common templates in the field, which are then compared and contrasted both within pools of similar methods and across lines of research. We provide a breakdown of text-to-image generation into various flavors of image-from-text methods, video-from-text methods, image editing, self-supervised and graph-based approaches. In this discussion, we focus on research papers published at 8 leading machine learning conferences in the years 2016-2022, also incorporating a number of relevant papers not matching the outlined search criteria. The conducted review suggests a significant increase in the number of papers published in the area and highlights research gaps and potential lines of investigation. To our knowledge, this is the first review to systematically look at text-to-image generation from the perspective of "cross-modal generation."
title Text-to-Image Cross-Modal Generation: A Systematic Review
topic Computer Vision and Pattern Recognition
Computation and Language
Machine Learning
url https://arxiv.org/abs/2401.11631