CulturalFrames: Assessing Cultural Expectation Alignment in Text-to-Image Models and Evaluation Metrics

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Nayak, Shravan, Bhatia, Mehar, Zhang, Xiaofeng, Rieser, Verena, Hendricks, Lisa Anne, van Steenkiste, Sjoerd, Goyal, Yash, Stańczak, Karolina, Agrawal, Aishwarya
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915736340922368
author Nayak, Shravan
Bhatia, Mehar
Zhang, Xiaofeng
Rieser, Verena
Hendricks, Lisa Anne
van Steenkiste, Sjoerd
Goyal, Yash
Stańczak, Karolina
Agrawal, Aishwarya
author_facet Nayak, Shravan
Bhatia, Mehar
Zhang, Xiaofeng
Rieser, Verena
Hendricks, Lisa Anne
van Steenkiste, Sjoerd
Goyal, Yash
Stańczak, Karolina
Agrawal, Aishwarya
contents The increasing ubiquity of text-to-image (T2I) models as tools for visual content generation raises concerns about their ability to accurately represent diverse cultural contexts -- where missed cues can stereotype communities and undermine usability. In this work, we present the first study to systematically quantify the alignment of T2I models and evaluation metrics with respect to both explicit (stated) as well as implicit (unstated, implied by the prompt's cultural context) cultural expectations. To this end, we introduce CulturalFrames, a novel benchmark designed for rigorous human evaluation of cultural representation in visual generations. Spanning 10 countries and 5 socio-cultural domains, CulturalFrames comprises 983 prompts, 3637 corresponding images generated by 4 state-of-the-art T2I models, and over 10k detailed human annotations. We find that across models and countries, cultural expectations are missed an average of 44% of the time. Among these failures, explicit expectations are missed at a surprisingly high average rate of 68%, while implicit expectation failures are also significant, averaging 49%. Furthermore, we show that existing T2I evaluation metrics correlate poorly with human judgments of cultural alignment, irrespective of their internal reasoning. Collectively, our findings expose critical gaps, provide a concrete testbed, and outline actionable directions for developing culturally informed T2I models and metrics that improve global usability.
format Preprint
id arxiv_https___arxiv_org_abs_2506_08835
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CulturalFrames: Assessing Cultural Expectation Alignment in Text-to-Image Models and Evaluation Metrics
Nayak, Shravan
Bhatia, Mehar
Zhang, Xiaofeng
Rieser, Verena
Hendricks, Lisa Anne
van Steenkiste, Sjoerd
Goyal, Yash
Stańczak, Karolina
Agrawal, Aishwarya
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
The increasing ubiquity of text-to-image (T2I) models as tools for visual content generation raises concerns about their ability to accurately represent diverse cultural contexts -- where missed cues can stereotype communities and undermine usability. In this work, we present the first study to systematically quantify the alignment of T2I models and evaluation metrics with respect to both explicit (stated) as well as implicit (unstated, implied by the prompt's cultural context) cultural expectations. To this end, we introduce CulturalFrames, a novel benchmark designed for rigorous human evaluation of cultural representation in visual generations. Spanning 10 countries and 5 socio-cultural domains, CulturalFrames comprises 983 prompts, 3637 corresponding images generated by 4 state-of-the-art T2I models, and over 10k detailed human annotations. We find that across models and countries, cultural expectations are missed an average of 44% of the time. Among these failures, explicit expectations are missed at a surprisingly high average rate of 68%, while implicit expectation failures are also significant, averaging 49%. Furthermore, we show that existing T2I evaluation metrics correlate poorly with human judgments of cultural alignment, irrespective of their internal reasoning. Collectively, our findings expose critical gaps, provide a concrete testbed, and outline actionable directions for developing culturally informed T2I models and metrics that improve global usability.
title CulturalFrames: Assessing Cultural Expectation Alignment in Text-to-Image Models and Evaluation Metrics
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2506.08835