CuRe: Cultural Gaps in the Long Tail of Text-to-Image Systems

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Rege, Aniket, Nie, Zinnia, Ramesh, Mahesh, Raskar, Unmesh, Yu, Zhuoran, Kusupati, Aditya, Lee, Yong Jae, Vinayak, Ramya Korlakai
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908401888395264
author Rege, Aniket
Nie, Zinnia
Ramesh, Mahesh
Raskar, Unmesh
Yu, Zhuoran
Kusupati, Aditya
Lee, Yong Jae
Vinayak, Ramya Korlakai
author_facet Rege, Aniket
Nie, Zinnia
Ramesh, Mahesh
Raskar, Unmesh
Yu, Zhuoran
Kusupati, Aditya
Lee, Yong Jae
Vinayak, Ramya Korlakai
contents Popular text-to-image (T2I) systems are trained on web-scraped data, which is heavily Amero and Euro-centric, underrepresenting the cultures of the Global South. To analyze these biases, we introduce CuRe, a novel and scalable benchmarking and scoring suite for cultural representativeness that leverages the marginal utility of attribute specification to T2I systems as a proxy for human judgments. Our CuRe benchmark dataset has a novel categorical hierarchy built from the crowdsourced Wikimedia knowledge graph, with 300 cultural artifacts across 32 cultural subcategories grouped into six broad cultural axes (food, art, fashion, architecture, celebrations, and people). Our dataset's categorical hierarchy enables CuRe scorers to evaluate T2I systems by analyzing their response to increasing the informativeness of text conditioning, enabling fine-grained cultural comparisons. We empirically observe much stronger correlations of our class of scorers to human judgments of perceptual similarity, image-text alignment, and cultural diversity across image encoders (SigLIP 2, AIMV2 and DINOv2), vision-language models (OpenCLIP, SigLIP 2, Gemini 2.0 Flash) and state-of-the-art text-to-image systems, including three variants of Stable Diffusion (1.5, XL, 3.5 Large), FLUX.1 [dev], Ideogram 2.0, and DALL-E 3. The code and dataset is open-sourced and available at https://aniketrege.github.io/cure/.
format Preprint
id arxiv_https___arxiv_org_abs_2506_08071
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CuRe: Cultural Gaps in the Long Tail of Text-to-Image Systems
Rege, Aniket
Nie, Zinnia
Ramesh, Mahesh
Raskar, Unmesh
Yu, Zhuoran
Kusupati, Aditya
Lee, Yong Jae
Vinayak, Ramya Korlakai
Computer Vision and Pattern Recognition
Popular text-to-image (T2I) systems are trained on web-scraped data, which is heavily Amero and Euro-centric, underrepresenting the cultures of the Global South. To analyze these biases, we introduce CuRe, a novel and scalable benchmarking and scoring suite for cultural representativeness that leverages the marginal utility of attribute specification to T2I systems as a proxy for human judgments. Our CuRe benchmark dataset has a novel categorical hierarchy built from the crowdsourced Wikimedia knowledge graph, with 300 cultural artifacts across 32 cultural subcategories grouped into six broad cultural axes (food, art, fashion, architecture, celebrations, and people). Our dataset's categorical hierarchy enables CuRe scorers to evaluate T2I systems by analyzing their response to increasing the informativeness of text conditioning, enabling fine-grained cultural comparisons. We empirically observe much stronger correlations of our class of scorers to human judgments of perceptual similarity, image-text alignment, and cultural diversity across image encoders (SigLIP 2, AIMV2 and DINOv2), vision-language models (OpenCLIP, SigLIP 2, Gemini 2.0 Flash) and state-of-the-art text-to-image systems, including three variants of Stable Diffusion (1.5, XL, 3.5 Large), FLUX.1 [dev], Ideogram 2.0, and DALL-E 3. The code and dataset is open-sourced and available at https://aniketrege.github.io/cure/.
title CuRe: Cultural Gaps in the Long Tail of Text-to-Image Systems
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.08071