Deconstructing Bias: A Multifaceted Framework for Diagnosing Cultural and Compositional Inequities in Text-to-Image Generative Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Said, Muna Numan, Zaidi, Aarib, Usman, Rabia, Okon, Sonia, Medepalli, Praneeth, Zhu, Kevin, Sharma, Vasu, O'Brien, Sean
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915271015399424
author Said, Muna Numan
Zaidi, Aarib
Usman, Rabia
Okon, Sonia
Medepalli, Praneeth
Zhu, Kevin
Sharma, Vasu
O'Brien, Sean
author_facet Said, Muna Numan
Zaidi, Aarib
Usman, Rabia
Okon, Sonia
Medepalli, Praneeth
Zhu, Kevin
Sharma, Vasu
O'Brien, Sean
contents The transformative potential of text-to-image (T2I) models hinges on their ability to synthesize culturally diverse, photorealistic images from textual prompts. However, these models often perpetuate cultural biases embedded within their training data, leading to systemic misrepresentations. This paper benchmarks the Component Inclusion Score (CIS), a metric designed to evaluate the fidelity of image generation across cultural contexts. Through extensive analysis involving 2,400 images, we quantify biases in terms of compositional fragility and contextual misalignment, revealing significant performance gaps between Western and non-Western cultural prompts. Our findings underscore the impact of data imbalance, attention entropy, and embedding superposition on model fairness. By benchmarking models like Stable Diffusion with CIS, we provide insights into architectural and data-centric interventions for enhancing cultural inclusivity in AI-generated imagery. This work advances the field by offering a comprehensive tool for diagnosing and mitigating biases in T2I generation, advocating for more equitable AI systems.
format Preprint
id arxiv_https___arxiv_org_abs_2505_01430
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Deconstructing Bias: A Multifaceted Framework for Diagnosing Cultural and Compositional Inequities in Text-to-Image Generative Models
Said, Muna Numan
Zaidi, Aarib
Usman, Rabia
Okon, Sonia
Medepalli, Praneeth
Zhu, Kevin
Sharma, Vasu
O'Brien, Sean
Computer Vision and Pattern Recognition
The transformative potential of text-to-image (T2I) models hinges on their ability to synthesize culturally diverse, photorealistic images from textual prompts. However, these models often perpetuate cultural biases embedded within their training data, leading to systemic misrepresentations. This paper benchmarks the Component Inclusion Score (CIS), a metric designed to evaluate the fidelity of image generation across cultural contexts. Through extensive analysis involving 2,400 images, we quantify biases in terms of compositional fragility and contextual misalignment, revealing significant performance gaps between Western and non-Western cultural prompts. Our findings underscore the impact of data imbalance, attention entropy, and embedding superposition on model fairness. By benchmarking models like Stable Diffusion with CIS, we provide insights into architectural and data-centric interventions for enhancing cultural inclusivity in AI-generated imagery. This work advances the field by offering a comprehensive tool for diagnosing and mitigating biases in T2I generation, advocating for more equitable AI systems.
title Deconstructing Bias: A Multifaceted Framework for Diagnosing Cultural and Compositional Inequities in Text-to-Image Generative Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.01430