All Seeds Are Not Equal: Enhancing Compositional Text-to-Image Generation with Reliable Random Seeds

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Li, Shuangqi, Le, Hieu, Xu, Jingyi, Salzmann, Mathieu
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866913746746605568
author Li, Shuangqi
Le, Hieu
Xu, Jingyi
Salzmann, Mathieu
author_facet Li, Shuangqi
Le, Hieu
Xu, Jingyi
Salzmann, Mathieu
contents Text-to-image diffusion models have demonstrated remarkable capability in generating realistic images from arbitrary text prompts. However, they often produce inconsistent results for compositional prompts such as "two dogs" or "a penguin on the right of a bowl". Understanding these inconsistencies is crucial for reliable image generation. In this paper, we highlight the significant role of initial noise in these inconsistencies, where certain noise patterns are more reliable for compositional prompts than others. Our analyses reveal that different initial random seeds tend to guide the model to place objects in distinct image areas, potentially adhering to specific patterns of camera angles and image composition associated with the seed. To improve the model's compositional ability, we propose a method for mining these reliable cases, resulting in a curated training set of generated images without requiring any manual annotation. By fine-tuning text-to-image models on these generated images, we significantly enhance their compositional capabilities. For numerical composition, we observe relative increases of 29.3% and 19.5% for Stable Diffusion and PixArt-α, respectively. Spatial composition sees even larger gains, with 60.7% for Stable Diffusion and 21.1% for PixArt-α.
format Preprint
id arxiv_https___arxiv_org_abs_2411_18810
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle All Seeds Are Not Equal: Enhancing Compositional Text-to-Image Generation with Reliable Random Seeds
Li, Shuangqi
Le, Hieu
Xu, Jingyi
Salzmann, Mathieu
Computer Vision and Pattern Recognition
Machine Learning
Text-to-image diffusion models have demonstrated remarkable capability in generating realistic images from arbitrary text prompts. However, they often produce inconsistent results for compositional prompts such as "two dogs" or "a penguin on the right of a bowl". Understanding these inconsistencies is crucial for reliable image generation. In this paper, we highlight the significant role of initial noise in these inconsistencies, where certain noise patterns are more reliable for compositional prompts than others. Our analyses reveal that different initial random seeds tend to guide the model to place objects in distinct image areas, potentially adhering to specific patterns of camera angles and image composition associated with the seed. To improve the model's compositional ability, we propose a method for mining these reliable cases, resulting in a curated training set of generated images without requiring any manual annotation. By fine-tuning text-to-image models on these generated images, we significantly enhance their compositional capabilities. For numerical composition, we observe relative increases of 29.3% and 19.5% for Stable Diffusion and PixArt-α, respectively. Spatial composition sees even larger gains, with 60.7% for Stable Diffusion and 21.1% for PixArt-α.
title All Seeds Are Not Equal: Enhancing Compositional Text-to-Image Generation with Reliable Random Seeds
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2411.18810