Entangled in Representations: Mechanistic Investigation of Cultural Biases in Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Yu, Haeun, Jeong, Seogyeong, Pawar, Siddhesh, Shin, Jisu, Jin, Jiho, Myung, Junho, Oh, Alice, Augenstein, Isabelle |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Survey of Cultural Awareness in Language Models: Text and Beyond
by: Pawar, Siddhesh, et al.
Published: (2024)
by: Pawar, Siddhesh, et al.
Published: (2024)
BiasGym: A Simple and Generalizable Framework for Analyzing and Removing Biases through Elicitation
by: Islam, Sekh Mainul, et al.
Published: (2025)
by: Islam, Sekh Mainul, et al.
Published: (2025)
Not What, But How: A Communicative Audit of LLM Response Framing
by: Pawar, Siddhesh Milind, et al.
Published: (2026)
by: Pawar, Siddhesh Milind, et al.
Published: (2026)
Presumed Cultural Identity: How Names Shape LLM Responses
by: Pawar, Siddhesh, et al.
Published: (2025)
by: Pawar, Siddhesh, et al.
Published: (2025)
MUG-Eval: A Proxy Evaluation Framework for Multilingual Generation Capabilities in Any Language
by: Song, Seyoung, et al.
Published: (2025)
by: Song, Seyoung, et al.
Published: (2025)
Social Bias Benchmark for Generation: A Comparison of Generation and QA-Based Evaluations
by: Jin, Jiho, et al.
Published: (2025)
by: Jin, Jiho, et al.
Published: (2025)
Revealing the Parametric Knowledge of Language Models: A Unified Framework for Attribution Methods
by: Yu, Haeun, et al.
Published: (2024)
by: Yu, Haeun, et al.
Published: (2024)
JuICE: A Benchmark for Evaluating LLM-Judge in Identifying Cultural Errors
by: Jin, Jiho, et al.
Published: (2026)
by: Jin, Jiho, et al.
Published: (2026)
LLM-C3MOD: A Human-LLM Collaborative System for Cross-Cultural Hate Speech Moderation
by: Park, Junyeong, et al.
Published: (2025)
by: Park, Junyeong, et al.
Published: (2025)
Exploring Cross-Cultural Differences in English Hate Speech Annotations: From Dataset Construction to Analysis
by: Lee, Nayeon, et al.
Published: (2023)
by: Lee, Nayeon, et al.
Published: (2023)
PapersPlease: A Benchmark for Evaluating Motivational Values of Large Language Models Based on ERG Theory
by: Myung, Junho, et al.
Published: (2025)
by: Myung, Junho, et al.
Published: (2025)
FINEST: Improving LLM Responses to Sensitive Topics Through Fine-Grained Evaluation
by: Oh, Juhyun, et al.
Published: (2026)
by: Oh, Juhyun, et al.
Published: (2026)
Quantifying Gender Biases Towards Politicians on Reddit
by: Marjanovic, Sara, et al.
Published: (2021)
by: Marjanovic, Sara, et al.
Published: (2021)
Code-Switching In-Context Learning for Cross-Lingual Transfer of Large Language Models
by: Yoo, Haneul, et al.
Published: (2025)
by: Yoo, Haneul, et al.
Published: (2025)
Probing Pre-Trained Language Models for Cross-Cultural Differences in Values
by: Arora, Arnav, et al.
Published: (2022)
by: Arora, Arnav, et al.
Published: (2022)
DYNAMICQA: Tracing Internal Knowledge Conflicts in Language Models
by: Marjanović, Sara Vera, et al.
Published: (2024)
by: Marjanović, Sara Vera, et al.
Published: (2024)
Understanding the Interplay between LLMs' Utilisation of Parametric and Contextual Knowledge: A keynote at ECIR 2025
by: Augenstein, Isabelle
Published: (2026)
by: Augenstein, Isabelle
Published: (2026)
CUB: Benchmarking Context Utilisation Techniques for Language Models
by: Hagström, Lovisa, et al.
Published: (2025)
by: Hagström, Lovisa, et al.
Published: (2025)
Measuring and Benchmarking Large Language Models' Capabilities to Generate Persuasive Language
by: Pauli, Amalie Brogaard, et al.
Published: (2024)
by: Pauli, Amalie Brogaard, et al.
Published: (2024)
LoCar: Localization-Aware Evaluation of In-Vehicle Assistants through Fine-Grained Sociolinguistic Control
by: Jeong, Seogyeong, et al.
Published: (2026)
by: Jeong, Seogyeong, et al.
Published: (2026)
Investigating the Impact of Model Instability on Explanations and Uncertainty
by: Marjanović, Sara Vera, et al.
Published: (2024)
by: Marjanović, Sara Vera, et al.
Published: (2024)
Spotting Out-of-Character Behavior: Atomic-Level Evaluation of Persona Fidelity in Open-Ended Generation
by: Shin, Jisu, et al.
Published: (2025)
by: Shin, Jisu, et al.
Published: (2025)
Aggregating Soft Labels from Crowd Annotations Improves Uncertainty Estimation Under Distribution Shift
by: Wright, Dustin, et al.
Published: (2022)
by: Wright, Dustin, et al.
Published: (2022)
Mitigating Cross-Lingual Cultural Inconsistencies in LLMs via Consensus-Driven Preference Optimisation
by: Resck, Lucas, et al.
Published: (2026)
by: Resck, Lucas, et al.
Published: (2026)
Perceptions to Beliefs: Exploring Precursory Inferences for Theory of Mind in Large Language Models
by: Jung, Chani, et al.
Published: (2024)
by: Jung, Chani, et al.
Published: (2024)
Claim Verification in the Age of Large Language Models: A Survey
by: Dmonte, Alphaeus, et al.
Published: (2024)
by: Dmonte, Alphaeus, et al.
Published: (2024)
Shared Heritage, Distinct Writing: Rethinking Resource Selection for East Asian Historical Documents
by: Song, Seyoung, et al.
Published: (2024)
by: Song, Seyoung, et al.
Published: (2024)
HERITAGE: An End-to-End Web Platform for Processing Korean Historical Documents in Hanja
by: Song, Seyoung, et al.
Published: (2025)
by: Song, Seyoung, et al.
Published: (2025)
A Reality Check on Context Utilisation for Retrieval-Augmented Generation
by: Hagström, Lovisa, et al.
Published: (2024)
by: Hagström, Lovisa, et al.
Published: (2024)
Ask LLMs Directly, "What shapes your bias?": Measuring Social Bias in Large Language Models
by: Shin, Jisu, et al.
Published: (2024)
by: Shin, Jisu, et al.
Published: (2024)
MentalBench: A DSM-Grounded Benchmark for Evaluating Psychiatric Diagnostic Capability of Large Language Models
by: Song, Hoyun, et al.
Published: (2026)
by: Song, Hoyun, et al.
Published: (2026)
Specializing Large Language Models to Simulate Survey Response Distributions for Global Populations
by: Cao, Yong, et al.
Published: (2025)
by: Cao, Yong, et al.
Published: (2025)
Language Models Entangle Language and Culture
by: Jain, Shourya, et al.
Published: (2026)
by: Jain, Shourya, et al.
Published: (2026)
Expanding Computation Spaces of LLMs at Inference Time
by: Jang, Yoonna, et al.
Published: (2025)
by: Jang, Yoonna, et al.
Published: (2025)
Social Bias Probing: Fairness Benchmarking for Language Models
by: Manerba, Marta Marchiori, et al.
Published: (2023)
by: Manerba, Marta Marchiori, et al.
Published: (2023)
RoleConflictBench: A Benchmark of Role Conflict Scenarios for Evaluating LLMs' Contextual Sensitivity
by: Shin, Jisu, et al.
Published: (2025)
by: Shin, Jisu, et al.
Published: (2025)
Cultural Biases of Large Language Models and Humans in Historical Interpretation
by: Celli, Fabio, et al.
Published: (2025)
by: Celli, Fabio, et al.
Published: (2025)
From Bytes to Biases: Investigating the Cultural Self-Perception of Large Language Models
by: Messner, Wolfgang, et al.
Published: (2023)
by: Messner, Wolfgang, et al.
Published: (2023)
Doc-PP: Document Policy Preservation Benchmark for Large Vision-Language Models
by: Jang, Haeun, et al.
Published: (2026)
by: Jang, Haeun, et al.
Published: (2026)
Revealing Fine-Grained Values and Opinions in Large Language Models
by: Wright, Dustin, et al.
Published: (2024)
by: Wright, Dustin, et al.
Published: (2024)
Similar Items
-
Survey of Cultural Awareness in Language Models: Text and Beyond
by: Pawar, Siddhesh, et al.
Published: (2024) -
BiasGym: A Simple and Generalizable Framework for Analyzing and Removing Biases through Elicitation
by: Islam, Sekh Mainul, et al.
Published: (2025) -
Not What, But How: A Communicative Audit of LLM Response Framing
by: Pawar, Siddhesh Milind, et al.
Published: (2026) -
Presumed Cultural Identity: How Names Shape LLM Responses
by: Pawar, Siddhesh, et al.
Published: (2025) -
MUG-Eval: A Proxy Evaluation Framework for Multilingual Generation Capabilities in Any Language
by: Song, Seyoung, et al.
Published: (2025)