Scaling Cultural Resources for Improving Generative Models
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912675970154496 |
|---|---|
| author | Stepanyan, Hayk Verma, Aishwarya Zaldivar, Andrew Feman, Rutledge Chin van Liemt, Erin MacMurray Kalia, Charu Prabhakaran, Vinodkumar Dev, Sunipa |
| author_facet | Stepanyan, Hayk Verma, Aishwarya Zaldivar, Andrew Feman, Rutledge Chin van Liemt, Erin MacMurray Kalia, Charu Prabhakaran, Vinodkumar Dev, Sunipa |
| contents | Generative models are known to have reduced performance in different global cultural contexts and languages. While continual data updates have been commonly conducted to improve overall model performance, bolstering and evaluating this cross-cultural competence of generative AI models requires data resources to be intentionally expanded to include global contexts and languages. In this work, we construct a repeatable, scalable, multi-pronged pipeline to collect and contribute culturally salient, multilingual data. We posit that such data can assess the state of the global applicability of our models and thus, in turn, help identify and improve upon cross-cultural gaps. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2510_25167 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Scaling Cultural Resources for Improving Generative Models Stepanyan, Hayk Verma, Aishwarya Zaldivar, Andrew Feman, Rutledge Chin van Liemt, Erin MacMurray Kalia, Charu Prabhakaran, Vinodkumar Dev, Sunipa Computers and Society Generative models are known to have reduced performance in different global cultural contexts and languages. While continual data updates have been commonly conducted to improve overall model performance, bolstering and evaluating this cross-cultural competence of generative AI models requires data resources to be intentionally expanded to include global contexts and languages. In this work, we construct a repeatable, scalable, multi-pronged pipeline to collect and contribute culturally salient, multilingual data. We posit that such data can assess the state of the global applicability of our models and thus, in turn, help identify and improve upon cross-cultural gaps. |
| title | Scaling Cultural Resources for Improving Generative Models |
| topic | Computers and Society |
| url | https://arxiv.org/abs/2510.25167 |