Scaling Cultural Resources for Improving Generative Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Stepanyan, Hayk, Verma, Aishwarya, Zaldivar, Andrew, Feman, Rutledge Chin, van Liemt, Erin MacMurray, Kalia, Charu, Prabhakaran, Vinodkumar, Dev, Sunipa
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912675970154496
author Stepanyan, Hayk
Verma, Aishwarya
Zaldivar, Andrew
Feman, Rutledge Chin
van Liemt, Erin MacMurray
Kalia, Charu
Prabhakaran, Vinodkumar
Dev, Sunipa
author_facet Stepanyan, Hayk
Verma, Aishwarya
Zaldivar, Andrew
Feman, Rutledge Chin
van Liemt, Erin MacMurray
Kalia, Charu
Prabhakaran, Vinodkumar
Dev, Sunipa
contents Generative models are known to have reduced performance in different global cultural contexts and languages. While continual data updates have been commonly conducted to improve overall model performance, bolstering and evaluating this cross-cultural competence of generative AI models requires data resources to be intentionally expanded to include global contexts and languages. In this work, we construct a repeatable, scalable, multi-pronged pipeline to collect and contribute culturally salient, multilingual data. We posit that such data can assess the state of the global applicability of our models and thus, in turn, help identify and improve upon cross-cultural gaps.
format Preprint
id arxiv_https___arxiv_org_abs_2510_25167
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Scaling Cultural Resources for Improving Generative Models
Stepanyan, Hayk
Verma, Aishwarya
Zaldivar, Andrew
Feman, Rutledge Chin
van Liemt, Erin MacMurray
Kalia, Charu
Prabhakaran, Vinodkumar
Dev, Sunipa
Computers and Society
Generative models are known to have reduced performance in different global cultural contexts and languages. While continual data updates have been commonly conducted to improve overall model performance, bolstering and evaluating this cross-cultural competence of generative AI models requires data resources to be intentionally expanded to include global contexts and languages. In this work, we construct a repeatable, scalable, multi-pronged pipeline to collect and contribute culturally salient, multilingual data. We posit that such data can assess the state of the global applicability of our models and thus, in turn, help identify and improve upon cross-cultural gaps.
title Scaling Cultural Resources for Improving Generative Models
topic Computers and Society
url https://arxiv.org/abs/2510.25167