Evaluating AI-Generated Images of Cultural Artifacts with Community-Informed Rubrics

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Johnson, Nari, Sudharsan, Deepthi, Hamna, Dalal, Samantha, Holroyd, Theo, Thieme, Anja, Heidari, Hoda, Massiceti, Daniela, Vaughan, Jennifer Wortman, Morrison, Cecily
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866914570934681600
author Johnson, Nari
Sudharsan, Deepthi
Hamna
Dalal, Samantha
Holroyd, Theo
Thieme, Anja
Heidari, Hoda
Massiceti, Daniela
Vaughan, Jennifer Wortman
Morrison, Cecily
author_facet Johnson, Nari
Sudharsan, Deepthi
Hamna
Dalal, Samantha
Holroyd, Theo
Thieme, Anja
Heidari, Hoda
Massiceti, Daniela
Vaughan, Jennifer Wortman
Morrison, Cecily
contents Measurement is essential to improving AI performance and mitigating harms for marginalized groups. As generative AI systems are rapidly deployed across geographies and contexts, AI measurement practices must be designed to support repeatable, automatable application across different models, datasets, and evaluation settings. But the drive to automate measurement can be in tension with the ability for measurement instruments to capture the expertise and perspectives of communities impacted by AI. Recent work advocates for breaking measurement into several key stages: first moving from an abstract concept to be measured into a precise, "systematized" concept; next operationalizing the systematized concept into a concrete measurement instrument; and finally applying the measurement instrument on data to produce measurements. This opens up an opportunity to concentrate community engagement in the systematization phase before operationalizing and applying measurement instruments. In this paper, we explore how to involve communities in systematizing the concept of "cultural appropriateness" in text-to-image models' representation of culturally significant artifacts through case studies with three communities: blind and low vision individuals residing in the UK, residents of Kerala, and residents of Tamil Nadu. Our systematized concepts reflect community members' lived experiences interacting with each artifact and how they want their material culture to be depicted, demonstrating the value of community involvement in defining valid measures. We explore how these systematized concepts can be operationalized into automated measurement instruments that could be applied using a multimodal LLM-as-a-judge approach and challenges that remain. We reflect on the benefits and limitations of such approaches.
format Preprint
id arxiv_https___arxiv_org_abs_2604_02406
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Evaluating AI-Generated Images of Cultural Artifacts with Community-Informed Rubrics
Johnson, Nari
Sudharsan, Deepthi
Hamna
Dalal, Samantha
Holroyd, Theo
Thieme, Anja
Heidari, Hoda
Massiceti, Daniela
Vaughan, Jennifer Wortman
Morrison, Cecily
Computers and Society
Measurement is essential to improving AI performance and mitigating harms for marginalized groups. As generative AI systems are rapidly deployed across geographies and contexts, AI measurement practices must be designed to support repeatable, automatable application across different models, datasets, and evaluation settings. But the drive to automate measurement can be in tension with the ability for measurement instruments to capture the expertise and perspectives of communities impacted by AI. Recent work advocates for breaking measurement into several key stages: first moving from an abstract concept to be measured into a precise, "systematized" concept; next operationalizing the systematized concept into a concrete measurement instrument; and finally applying the measurement instrument on data to produce measurements. This opens up an opportunity to concentrate community engagement in the systematization phase before operationalizing and applying measurement instruments. In this paper, we explore how to involve communities in systematizing the concept of "cultural appropriateness" in text-to-image models' representation of culturally significant artifacts through case studies with three communities: blind and low vision individuals residing in the UK, residents of Kerala, and residents of Tamil Nadu. Our systematized concepts reflect community members' lived experiences interacting with each artifact and how they want their material culture to be depicted, demonstrating the value of community involvement in defining valid measures. We explore how these systematized concepts can be operationalized into automated measurement instruments that could be applied using a multimodal LLM-as-a-judge approach and challenges that remain. We reflect on the benefits and limitations of such approaches.
title Evaluating AI-Generated Images of Cultural Artifacts with Community-Informed Rubrics
topic Computers and Society
url https://arxiv.org/abs/2604.02406