Reproducibility Study of "ITI-GEN: Inclusive Text-to-Image Generation"

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Fernández, Daniel Gallo, Matisan, Răzvan-Andrei, Muñoz, Alejandro Monroy, Partyka, Janusz
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909272656314368
author Fernández, Daniel Gallo
Matisan, Răzvan-Andrei
Muñoz, Alejandro Monroy
Partyka, Janusz
author_facet Fernández, Daniel Gallo
Matisan, Răzvan-Andrei
Muñoz, Alejandro Monroy
Partyka, Janusz
contents Text-to-image generative models often present issues regarding fairness with respect to certain sensitive attributes, such as gender or skin tone. This study aims to reproduce the results presented in "ITI-GEN: Inclusive Text-to-Image Generation" by Zhang et al. (2023a), which introduces a model to improve inclusiveness in these kinds of models. We show that most of the claims made by the authors about ITI-GEN hold: it improves the diversity and quality of generated images, it is scalable to different domains, it has plug-and-play capabilities, and it is efficient from a computational point of view. However, ITI-GEN sometimes uses undesired attributes as proxy features and it is unable to disentangle some pairs of (correlated) attributes such as gender and baldness. In addition, when the number of considered attributes increases, the training time grows exponentially and ITI-GEN struggles to generate inclusive images for all elements in the joint distribution. To solve these issues, we propose using Hard Prompt Search with negative prompting, a method that does not require training and that handles negation better than vanilla Hard Prompt Search. Nonetheless, Hard Prompt Search (with or without negative prompting) cannot be used for continuous attributes that are hard to express in natural language, an area where ITI-GEN excels as it is guided by images during training. Finally, we propose combining ITI-GEN and Hard Prompt Search with negative prompting.
format Preprint
id arxiv_https___arxiv_org_abs_2407_19996
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Reproducibility Study of "ITI-GEN: Inclusive Text-to-Image Generation"
Fernández, Daniel Gallo
Matisan, Răzvan-Andrei
Muñoz, Alejandro Monroy
Partyka, Janusz
Computer Vision and Pattern Recognition
Artificial Intelligence
Text-to-image generative models often present issues regarding fairness with respect to certain sensitive attributes, such as gender or skin tone. This study aims to reproduce the results presented in "ITI-GEN: Inclusive Text-to-Image Generation" by Zhang et al. (2023a), which introduces a model to improve inclusiveness in these kinds of models. We show that most of the claims made by the authors about ITI-GEN hold: it improves the diversity and quality of generated images, it is scalable to different domains, it has plug-and-play capabilities, and it is efficient from a computational point of view. However, ITI-GEN sometimes uses undesired attributes as proxy features and it is unable to disentangle some pairs of (correlated) attributes such as gender and baldness. In addition, when the number of considered attributes increases, the training time grows exponentially and ITI-GEN struggles to generate inclusive images for all elements in the joint distribution. To solve these issues, we propose using Hard Prompt Search with negative prompting, a method that does not require training and that handles negation better than vanilla Hard Prompt Search. Nonetheless, Hard Prompt Search (with or without negative prompting) cannot be used for continuous attributes that are hard to express in natural language, an area where ITI-GEN excels as it is guided by images during training. Finally, we propose combining ITI-GEN and Hard Prompt Search with negative prompting.
title Reproducibility Study of "ITI-GEN: Inclusive Text-to-Image Generation"
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2407.19996