Fluent but Foreign: Even Regional LLMs Lack Cultural Alignment

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Agarwal, Dhruv, Shukla, Anya, Sitaram, Sunayana, Vashistha, Aditya
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917218505195520
author Agarwal, Dhruv
Shukla, Anya
Sitaram, Sunayana
Vashistha, Aditya
author_facet Agarwal, Dhruv
Shukla, Anya
Sitaram, Sunayana
Vashistha, Aditya
contents Large language models (LLMs) are used worldwide, yet exhibit Western cultural tendencies. Many countries are now building ``regional'' or ``sovereign'' LLMs, but it remains unclear whether they reflect local values and practices or merely speak local languages. Using India as a case study, we evaluate six Indic and six global LLMs on two dimensions -- values and practices -- grounded in nationally representative surveys and community-sourced QA datasets. Across tasks, Indic models do not align better with Indian norms than global models; in fact, a U.S. respondent is a closer proxy for Indian values than any Indic model. We further run a user study with 115 Indian users and find that writing suggestions from both global and Indic LLMs introduce Westernized or exoticized writing. Prompting and regional fine-tuning fail to recover alignment and can even degrade existing knowledge. We attribute this to scarce culturally grounded data, especially for pretraining. We position cultural evaluation as a first-class requirement alongside multilingual benchmarks and offer a reusable, community-grounded methodology. We call for native, community-authored corpora and thickxwide evaluations to build truly sovereign LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2505_21548
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Fluent but Foreign: Even Regional LLMs Lack Cultural Alignment
Agarwal, Dhruv
Shukla, Anya
Sitaram, Sunayana
Vashistha, Aditya
Computation and Language
Artificial Intelligence
Computers and Society
Physics and Society
Large language models (LLMs) are used worldwide, yet exhibit Western cultural tendencies. Many countries are now building ``regional'' or ``sovereign'' LLMs, but it remains unclear whether they reflect local values and practices or merely speak local languages. Using India as a case study, we evaluate six Indic and six global LLMs on two dimensions -- values and practices -- grounded in nationally representative surveys and community-sourced QA datasets. Across tasks, Indic models do not align better with Indian norms than global models; in fact, a U.S. respondent is a closer proxy for Indian values than any Indic model. We further run a user study with 115 Indian users and find that writing suggestions from both global and Indic LLMs introduce Westernized or exoticized writing. Prompting and regional fine-tuning fail to recover alignment and can even degrade existing knowledge. We attribute this to scarce culturally grounded data, especially for pretraining. We position cultural evaluation as a first-class requirement alongside multilingual benchmarks and offer a reusable, community-grounded methodology. We call for native, community-authored corpora and thickxwide evaluations to build truly sovereign LLMs.
title Fluent but Foreign: Even Regional LLMs Lack Cultural Alignment
topic Computation and Language
Artificial Intelligence
Computers and Society
Physics and Society
url https://arxiv.org/abs/2505.21548