Improving face generation quality and prompt following with synthetic captions

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tarasiou, Michail, Moschoglou, Stylianos, Deng, Jiankang, Zafeiriou, Stefanos
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916250424180736
author Tarasiou, Michail
Moschoglou, Stylianos
Deng, Jiankang
Zafeiriou, Stefanos
author_facet Tarasiou, Michail
Moschoglou, Stylianos
Deng, Jiankang
Zafeiriou, Stefanos
contents Recent advancements in text-to-image generation using diffusion models have significantly improved the quality of generated images and expanded the ability to depict a wide range of objects. However, ensuring that these models adhere closely to the text prompts remains a considerable challenge. This issue is particularly pronounced when trying to generate photorealistic images of humans. Without significant prompt engineering efforts models often produce unrealistic images and typically fail to incorporate the full extent of the prompt information. This limitation can be largely attributed to the nature of captions accompanying the images used in training large scale diffusion models, which typically prioritize contextual information over details related to the person's appearance. In this paper we address this issue by introducing a training-free pipeline designed to generate accurate appearance descriptions from images of people. We apply this method to create approximately 250,000 captions for publicly available face datasets. We then use these synthetic captions to fine-tune a text-to-image diffusion model. Our results demonstrate that this approach significantly improves the model's ability to generate high-quality, realistic human faces and enhances adherence to the given prompts, compared to the baseline model. We share our synthetic captions, pretrained checkpoints and training code.
format Preprint
id arxiv_https___arxiv_org_abs_2405_10864
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Improving face generation quality and prompt following with synthetic captions
Tarasiou, Michail
Moschoglou, Stylianos
Deng, Jiankang
Zafeiriou, Stefanos
Computer Vision and Pattern Recognition
Machine Learning
Recent advancements in text-to-image generation using diffusion models have significantly improved the quality of generated images and expanded the ability to depict a wide range of objects. However, ensuring that these models adhere closely to the text prompts remains a considerable challenge. This issue is particularly pronounced when trying to generate photorealistic images of humans. Without significant prompt engineering efforts models often produce unrealistic images and typically fail to incorporate the full extent of the prompt information. This limitation can be largely attributed to the nature of captions accompanying the images used in training large scale diffusion models, which typically prioritize contextual information over details related to the person's appearance. In this paper we address this issue by introducing a training-free pipeline designed to generate accurate appearance descriptions from images of people. We apply this method to create approximately 250,000 captions for publicly available face datasets. We then use these synthetic captions to fine-tune a text-to-image diffusion model. Our results demonstrate that this approach significantly improves the model's ability to generate high-quality, realistic human faces and enhances adherence to the given prompts, compared to the baseline model. We share our synthetic captions, pretrained checkpoints and training code.
title Improving face generation quality and prompt following with synthetic captions
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2405.10864