Culture is Everywhere: A Call for Intentionally Cultural Evaluation
Fuente:
arXiv
Guardado en:
| Autores principales: | Oh, Juhyun, Cha, Inha, Saxon, Michael, Lim, Hyunseung, Bhatt, Shaily, Oh, Alice |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
The Generative AI Paradox on Evaluation: What It Can Solve, It May Not Evaluate
por: Oh, Juhyun, et al.
Publicado: (2024)
por: Oh, Juhyun, et al.
Publicado: (2024)
Extrinsic Evaluation of Cultural Competence in Large Language Models
por: Bhatt, Shaily, et al.
Publicado: (2024)
por: Bhatt, Shaily, et al.
Publicado: (2024)
Uncovering Factor Level Preferences to Improve Human-Model Alignment
por: Oh, Juhyun, et al.
Publicado: (2024)
por: Oh, Juhyun, et al.
Publicado: (2024)
Flex-TravelPlanner: A Benchmark for Flexible Planning with Language Agents
por: Oh, Juhyun, et al.
Publicado: (2025)
por: Oh, Juhyun, et al.
Publicado: (2025)
JuICE: A Benchmark for Evaluating LLM-Judge in Identifying Cultural Errors
por: Jin, Jiho, et al.
Publicado: (2026)
por: Jin, Jiho, et al.
Publicado: (2026)
Research Borderlands: Analysing Writing Across Research Cultures
por: Bhatt, Shaily, et al.
Publicado: (2025)
por: Bhatt, Shaily, et al.
Publicado: (2025)
Multi-FAct: Assessing Factuality of Multilingual LLMs using FActScore
por: Shafayat, Sheikh, et al.
Publicado: (2024)
por: Shafayat, Sheikh, et al.
Publicado: (2024)
OLA: Output Language Alignment in Code-Switched LLM Interactions
por: Oh, Juhyun, et al.
Publicado: (2026)
por: Oh, Juhyun, et al.
Publicado: (2026)
Spotting Out-of-Character Behavior: Atomic-Level Evaluation of Persona Fidelity in Open-Ended Generation
por: Shin, Jisu, et al.
Publicado: (2025)
por: Shin, Jisu, et al.
Publicado: (2025)
RoleConflictBench: A Benchmark of Role Conflict Scenarios for Evaluating LLMs' Contextual Sensitivity
por: Shin, Jisu, et al.
Publicado: (2025)
por: Shin, Jisu, et al.
Publicado: (2025)
CLIcK: A Benchmark Dataset of Cultural and Linguistic Intelligence in Korean
por: Kim, Eunsu, et al.
Publicado: (2024)
por: Kim, Eunsu, et al.
Publicado: (2024)
CAIRe: Cultural Attribution of Images by Retrieval-Augmented Evaluation
por: Yayavaram, Arnav, et al.
Publicado: (2025)
por: Yayavaram, Arnav, et al.
Publicado: (2025)
FINEST: Improving LLM Responses to Sensitive Topics Through Fine-Grained Evaluation
por: Oh, Juhyun, et al.
Publicado: (2026)
por: Oh, Juhyun, et al.
Publicado: (2026)
Are they lovers or friends? Evaluating LLMs' Social Reasoning in English and Korean Dialogues
por: Kim, Eunsu, et al.
Publicado: (2025)
por: Kim, Eunsu, et al.
Publicado: (2025)
TALES: A Taxonomy and Analysis of Cultural Representations in LLM-generated Stories
por: Bhagat, Kirti, et al.
Publicado: (2025)
por: Bhagat, Kirti, et al.
Publicado: (2025)
When Tom Eats Kimchi: Evaluating Cultural Bias of Multimodal Large Language Models in Cultural Mixture Contexts
por: Kim, Jun Seong, et al.
Publicado: (2025)
por: Kim, Jun Seong, et al.
Publicado: (2025)
Can LLM Generate Culturally Relevant Commonsense QA Data? Case Study in Indonesian and Sundanese
por: Putri, Rifki Afina, et al.
Publicado: (2024)
por: Putri, Rifki Afina, et al.
Publicado: (2024)
Benchmarks as Microscopes: A Call for Model Metrology
por: Saxon, Michael, et al.
Publicado: (2024)
por: Saxon, Michael, et al.
Publicado: (2024)
LLM-C3MOD: A Human-LLM Collaborative System for Cross-Cultural Hate Speech Moderation
por: Park, Junyeong, et al.
Publicado: (2025)
por: Park, Junyeong, et al.
Publicado: (2025)
Do You Know About My Nation? Investigating Multilingual Language Models' Cultural Literacy Through Factual Knowledge
por: Tanwar, Eshaan, et al.
Publicado: (2025)
por: Tanwar, Eshaan, et al.
Publicado: (2025)
BLUCK: A Benchmark Dataset for Bengali Linguistic Understanding and Cultural Knowledge
por: Kabir, Daeen, et al.
Publicado: (2025)
por: Kabir, Daeen, et al.
Publicado: (2025)
LLM-as-a-tutor in EFL Writing Education: Focusing on Evaluation of Student-LLM Interaction
por: Han, Jieun, et al.
Publicado: (2023)
por: Han, Jieun, et al.
Publicado: (2023)
Entangled in Representations: Mechanistic Investigation of Cultural Biases in Large Language Models
por: Yu, Haeun, et al.
Publicado: (2025)
por: Yu, Haeun, et al.
Publicado: (2025)
Exploring Cross-Cultural Differences in English Hate Speech Annotations: From Dataset Construction to Analysis
por: Lee, Nayeon, et al.
Publicado: (2023)
por: Lee, Nayeon, et al.
Publicado: (2023)
Survey of Cultural Awareness in Language Models: Text and Beyond
por: Pawar, Siddhesh, et al.
Publicado: (2024)
por: Pawar, Siddhesh, et al.
Publicado: (2024)
Exploring Persona-dependent LLM Alignment for the Moral Machine Experiment
por: Kim, Jiseon, et al.
Publicado: (2025)
por: Kim, Jiseon, et al.
Publicado: (2025)
Cultural Compass: A Framework for Organizing Societal Norms to Detect Violations in Human-AI Conversations
por: Cheng, Myra, et al.
Publicado: (2026)
por: Cheng, Myra, et al.
Publicado: (2026)
Social Bias Benchmark for Generation: A Comparison of Generation and QA-Based Evaluations
por: Jin, Jiho, et al.
Publicado: (2025)
por: Jin, Jiho, et al.
Publicado: (2025)
Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual Evaluation
por: Singh, Shivalika, et al.
Publicado: (2024)
por: Singh, Shivalika, et al.
Publicado: (2024)
Who Speaks Matters: Analysing the Influence of the Speaker's Ethnicity on Hate Classification
por: Malik, Ananya, et al.
Publicado: (2024)
por: Malik, Ananya, et al.
Publicado: (2024)
PapersPlease: A Benchmark for Evaluating Motivational Values of Large Language Models Based on ERG Theory
por: Myung, Junho, et al.
Publicado: (2025)
por: Myung, Junho, et al.
Publicado: (2025)
A 2-step Framework for Automated Literary Translation Evaluation: Its Promises and Pitfalls
por: Shafayat, Sheikh, et al.
Publicado: (2024)
por: Shafayat, Sheikh, et al.
Publicado: (2024)
CultureForest: Understanding and Evaluating Cultural Norm Grounded Reasoning in LLMs
por: Ye, Yangfan, et al.
Publicado: (2026)
por: Ye, Yangfan, et al.
Publicado: (2026)
CURE: Cultural Understanding and Reasoning Evaluation - A Framework for "Thick" Culture Alignment Evaluation in LLMs
por: Vo, Truong, et al.
Publicado: (2025)
por: Vo, Truong, et al.
Publicado: (2025)
On the Effect of Uncertainty on Layer-wise Inference Dynamics
por: Kim, Sunwoo, et al.
Publicado: (2025)
por: Kim, Sunwoo, et al.
Publicado: (2025)
MentalBench: A DSM-Grounded Benchmark for Evaluating Psychiatric Diagnostic Capability of Large Language Models
por: Song, Hoyun, et al.
Publicado: (2026)
por: Song, Hoyun, et al.
Publicado: (2026)
XL-SafetyBench: A Country-Grounded Cross-Cultural Benchmark for LLM Safety and Cultural Sensitivity
por: Choi, Dasol, et al.
Publicado: (2026)
por: Choi, Dasol, et al.
Publicado: (2026)
Can LLMs Express Personality Across Cultures? Introducing CulturalPersonas for Evaluating Trait Alignment
por: Dey, Priyanka, et al.
Publicado: (2025)
por: Dey, Priyanka, et al.
Publicado: (2025)
KULTURE Bench: A Benchmark for Assessing Language Model in Korean Cultural Context
por: Wang, Xiaonan, et al.
Publicado: (2024)
por: Wang, Xiaonan, et al.
Publicado: (2024)
Pragmatics Meets Culture: Culturally-adapted Artwork Description Generation and Evaluation
por: Zhao, Lingjun, et al.
Publicado: (2026)
por: Zhao, Lingjun, et al.
Publicado: (2026)
Ejemplares similares
-
The Generative AI Paradox on Evaluation: What It Can Solve, It May Not Evaluate
por: Oh, Juhyun, et al.
Publicado: (2024) -
Extrinsic Evaluation of Cultural Competence in Large Language Models
por: Bhatt, Shaily, et al.
Publicado: (2024) -
Uncovering Factor Level Preferences to Improve Human-Model Alignment
por: Oh, Juhyun, et al.
Publicado: (2024) -
Flex-TravelPlanner: A Benchmark for Flexible Planning with Language Agents
por: Oh, Juhyun, et al.
Publicado: (2025) -
JuICE: A Benchmark for Evaluating LLM-Judge in Identifying Cultural Errors
por: Jin, Jiho, et al.
Publicado: (2026)