How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Brack, Manuel, Katakol, Sudeep, Friedrich, Felix, Schramowski, Patrick, Ravi, Hareesh, Kersting, Kristian, Kale, Ajinkya
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908414753374208
author Brack, Manuel
Katakol, Sudeep
Friedrich, Felix
Schramowski, Patrick
Ravi, Hareesh
Kersting, Kristian
Kale, Ajinkya
author_facet Brack, Manuel
Katakol, Sudeep
Friedrich, Felix
Schramowski, Patrick
Ravi, Hareesh
Kersting, Kristian
Kale, Ajinkya
contents Training data is at the core of any successful text-to-image models. The quality and descriptiveness of image text are crucial to a model's performance. Given the noisiness and inconsistency in web-scraped datasets, recent works shifted towards synthetic training captions. While this setup is generally believed to produce more capable models, current literature does not provide any insights into its design choices. This study closes this gap by systematically investigating how different synthetic captioning strategies impact the downstream performance of text-to-image models. Our experiments demonstrate that dense, high-quality captions enhance text alignment but may introduce trade-offs in output aesthetics and diversity. Conversely, captions of randomized lengths yield balanced improvements across aesthetics and alignment without compromising sample diversity. We also demonstrate that varying caption distributions introduce significant shifts in the output bias of a trained model. Our findings underscore the importance of caption design in achieving optimal model performance and provide practical insights for more effective training data strategies in text-to-image generation.
format Preprint
id arxiv_https___arxiv_org_abs_2506_16679
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions
Brack, Manuel
Katakol, Sudeep
Friedrich, Felix
Schramowski, Patrick
Ravi, Hareesh
Kersting, Kristian
Kale, Ajinkya
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Training data is at the core of any successful text-to-image models. The quality and descriptiveness of image text are crucial to a model's performance. Given the noisiness and inconsistency in web-scraped datasets, recent works shifted towards synthetic training captions. While this setup is generally believed to produce more capable models, current literature does not provide any insights into its design choices. This study closes this gap by systematically investigating how different synthetic captioning strategies impact the downstream performance of text-to-image models. Our experiments demonstrate that dense, high-quality captions enhance text alignment but may introduce trade-offs in output aesthetics and diversity. Conversely, captions of randomized lengths yield balanced improvements across aesthetics and alignment without compromising sample diversity. We also demonstrate that varying caption distributions introduce significant shifts in the output bias of a trained model. Our findings underscore the importance of caption design in achieving optimal model performance and provide practical insights for more effective training data strategies in text-to-image generation.
title How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2506.16679