Entangled in Representations: Mechanistic Investigation of Cultural Biases in Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yu, Haeun, Jeong, Seogyeong, Pawar, Siddhesh, Shin, Jisu, Jin, Jiho, Myung, Junho, Oh, Alice, Augenstein, Isabelle
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914259195133952
author Yu, Haeun
Jeong, Seogyeong
Pawar, Siddhesh
Shin, Jisu
Jin, Jiho
Myung, Junho
Oh, Alice
Augenstein, Isabelle
author_facet Yu, Haeun
Jeong, Seogyeong
Pawar, Siddhesh
Shin, Jisu
Jin, Jiho
Myung, Junho
Oh, Alice
Augenstein, Isabelle
contents The growing deployment of large language models (LLMs) across diverse cultural contexts necessitates a deeper understanding of LLMs' representations of different cultures. Prior work has focused on evaluating the cultural awareness of LLMs by only examining the text they generate. This approach overlooks the internal sources of cultural misrepresentation within the models themselves. To bridge this gap, we propose Culturescope, the first mechanistic interpretability-based method that probes the internal representations of different cultural knowledge in LLMs. We also introduce a cultural flattening score as a measure of the intrinsic cultural biases of the decoded knowledge from Culturescope. Additionally, we study how LLMs internalize cultural biases, which allows us to trace how cultural biases such as Western-dominance bias and cultural flattening emerge within LLMs. We find that low-resource cultures are less susceptible to cultural biases, likely due to the model's limited parametric knowledge. Our work provides a foundation for future research on mitigating cultural biases and enhancing LLMs' cultural understanding.
format Preprint
id arxiv_https___arxiv_org_abs_2508_08879
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Entangled in Representations: Mechanistic Investigation of Cultural Biases in Large Language Models
Yu, Haeun
Jeong, Seogyeong
Pawar, Siddhesh
Shin, Jisu
Jin, Jiho
Myung, Junho
Oh, Alice
Augenstein, Isabelle
Computation and Language
Artificial Intelligence
The growing deployment of large language models (LLMs) across diverse cultural contexts necessitates a deeper understanding of LLMs' representations of different cultures. Prior work has focused on evaluating the cultural awareness of LLMs by only examining the text they generate. This approach overlooks the internal sources of cultural misrepresentation within the models themselves. To bridge this gap, we propose Culturescope, the first mechanistic interpretability-based method that probes the internal representations of different cultural knowledge in LLMs. We also introduce a cultural flattening score as a measure of the intrinsic cultural biases of the decoded knowledge from Culturescope. Additionally, we study how LLMs internalize cultural biases, which allows us to trace how cultural biases such as Western-dominance bias and cultural flattening emerge within LLMs. We find that low-resource cultures are less susceptible to cultural biases, likely due to the model's limited parametric knowledge. Our work provides a foundation for future research on mitigating cultural biases and enhancing LLMs' cultural understanding.
title Entangled in Representations: Mechanistic Investigation of Cultural Biases in Large Language Models
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2508.08879