CURE: Cultural Understanding and Reasoning Evaluation - A Framework for "Thick" Culture Alignment Evaluation in LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Vo, Truong, Koyejo, Sanmi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Toward Cultural Interpretability: A Linguistic Anthropological Framework for Describing and Evaluating Large Language Models (LLMs)
by: Jones, Graham M., et al.
Published: (2024)
by: Jones, Graham M., et al.
Published: (2024)
Pragmatics Meets Culture: Culturally-adapted Artwork Description Generation and Evaluation
by: Zhao, Lingjun, et al.
Published: (2026)
by: Zhao, Lingjun, et al.
Published: (2026)
Questionnaires for Everyone: Streamlining Cross-Cultural Questionnaire Adaptation with GPT-Based Translation Quality Evaluation
by: Haavisto, Otso, et al.
Published: (2024)
by: Haavisto, Otso, et al.
Published: (2024)
The Case for "Thick Evaluations" of Cultural Representation in AI
by: Qadri, Rida, et al.
Published: (2025)
by: Qadri, Rida, et al.
Published: (2025)
DaKultur: Evaluating the Cultural Awareness of Language Models for Danish with Native Speakers
by: Müller-Eberstein, Max, et al.
Published: (2025)
by: Müller-Eberstein, Max, et al.
Published: (2025)
CUPID: Evaluating Personalized and Contextualized Alignment of LLMs from Interactions
by: Kim, Tae Soo, et al.
Published: (2025)
by: Kim, Tae Soo, et al.
Published: (2025)
A Framework for Evaluating LLMs Under Task Indeterminacy
by: Guerdan, Luke, et al.
Published: (2024)
by: Guerdan, Luke, et al.
Published: (2024)
Robots in the Middle: Evaluating LLMs in Dispute Resolution
by: Tan, Jinzhe, et al.
Published: (2024)
by: Tan, Jinzhe, et al.
Published: (2024)
CHBench: A Cognitive Hierarchy Benchmark for Evaluating Strategic Reasoning Capability of LLMs
by: Liu, Hongtao, et al.
Published: (2025)
by: Liu, Hongtao, et al.
Published: (2025)
Beyond Screenshots: Evaluating VLMs' Understanding of UI Animations
by: Liang, Chen, et al.
Published: (2026)
by: Liang, Chen, et al.
Published: (2026)
DICE: A Framework for Dimensional and Contextual Evaluation of Language Models
by: Shrivastava, Aryan, et al.
Published: (2025)
by: Shrivastava, Aryan, et al.
Published: (2025)
LLAMADRS: Evaluating Open-Source LLMs on Real Clinical Interviews--To Reason or Not to Reason?
by: Kebe, Gaoussou Youssouf, et al.
Published: (2025)
by: Kebe, Gaoussou Youssouf, et al.
Published: (2025)
Alignment Drift in Multimodal LLMs: A Two-Phase, Longitudinal Evaluation of Harm Across Eight Model Releases
by: Ford, Casey, et al.
Published: (2026)
by: Ford, Casey, et al.
Published: (2026)
Unpacking Human Preference for LLMs: Demographically Aware Evaluation with the HUMAINE Framework
by: Petrova, Nora, et al.
Published: (2026)
by: Petrova, Nora, et al.
Published: (2026)
From Chatbots to Confidants: A Cross-Cultural Study of LLM Adoption for Emotional Support
by: Amat-Lefort, Natalia, et al.
Published: (2026)
by: Amat-Lefort, Natalia, et al.
Published: (2026)
Understanding Public Perceptions of AI Conversational Agents: A Cross-Cultural Analysis
by: Liu, Zihan, et al.
Published: (2024)
by: Liu, Zihan, et al.
Published: (2024)
Beyond Western Politics: Cross-Cultural Benchmarks for Evaluating Partisan Associations in LLMs
by: Kumar, Divyanshu, et al.
Published: (2025)
by: Kumar, Divyanshu, et al.
Published: (2025)
CulturalTeaming: AI-Assisted Interactive Red-Teaming for Challenging LLMs' (Lack of) Multicultural Knowledge
by: Chiu, Yu Ying, et al.
Published: (2024)
by: Chiu, Yu Ying, et al.
Published: (2024)
OPeRA: A Dataset of Observation, Persona, Rationale, and Action for Evaluating LLMs on Human Online Shopping Behavior Simulation
by: Wang, Ziyi, et al.
Published: (2025)
by: Wang, Ziyi, et al.
Published: (2025)
ValueCompass: A Framework for Measuring Contextual Value Alignment Between Human and LLMs
by: Shen, Hua, et al.
Published: (2024)
by: Shen, Hua, et al.
Published: (2024)
Coimagining the Future of Voice Assistants with Cultural Sensitivity
by: Seaborn, Katie, et al.
Published: (2024)
by: Seaborn, Katie, et al.
Published: (2024)
Evaluating LLMs as Human Surrogates in Controlled Experiments
by: Hoq, Adnan, et al.
Published: (2026)
by: Hoq, Adnan, et al.
Published: (2026)
Multi-Intent Recognition in Dialogue Understanding: A Comparison Between Smaller Open-Source LLMs
by: Ahmad, Adnan, et al.
Published: (2025)
by: Ahmad, Adnan, et al.
Published: (2025)
Culturally Grounded Personas in Large Language Models: Characterization and Alignment with Socio-Psychological Value Frameworks
by: Greco, Candida M., et al.
Published: (2026)
by: Greco, Candida M., et al.
Published: (2026)
Assessing LLMs in Art Contexts: Critique Generation and Theory of Mind Evaluation
by: Arita, Takaya, et al.
Published: (2025)
by: Arita, Takaya, et al.
Published: (2025)
Beyond Models: A Framework for Contextual and Cultural Intelligence in African AI Deployment
by: Ndlovu, Qness
Published: (2025)
by: Ndlovu, Qness
Published: (2025)
BEYOND DIALOGUE: A Profile-Dialogue Alignment Framework Towards General Role-Playing Language Model
by: Yu, Yeyong, et al.
Published: (2024)
by: Yu, Yeyong, et al.
Published: (2024)
Do LLMs Make Mistakes Like Students? Exploring Natural Alignment between Language Models and Human Error Patterns
by: Liu, Naiming, et al.
Published: (2025)
by: Liu, Naiming, et al.
Published: (2025)
A Scalable Framework for Evaluating Health Language Models
by: Mallinar, Neil, et al.
Published: (2025)
by: Mallinar, Neil, et al.
Published: (2025)
Value Alignment of Social Media Ranking Algorithms
by: Jahanbakhsh, Farnaz, et al.
Published: (2025)
by: Jahanbakhsh, Farnaz, et al.
Published: (2025)
Evaluating Behavioral Alignment in Conflict Dialogue: A Multi-Dimensional Comparison of LLM Agents and Humans
by: Kwon, Deuksin, et al.
Published: (2025)
by: Kwon, Deuksin, et al.
Published: (2025)
AIRepr: An Analyst-Inspector Framework for Evaluating Reproducibility of LLMs in Data Science
by: Zeng, Qiuhai, et al.
Published: (2025)
by: Zeng, Qiuhai, et al.
Published: (2025)
Advancing AI-Scientist Understanding: Multi-Agent LLMs with Interpretable Physics Reasoning
by: Xu, Yinggan, et al.
Published: (2025)
by: Xu, Yinggan, et al.
Published: (2025)
Meta-Evaluating Local LLMs: Rethinking Performance Metrics for Serious Games
by: Isaza-Giraldo, Andrés, et al.
Published: (2025)
by: Isaza-Giraldo, Andrés, et al.
Published: (2025)
Pearmut: Human Evaluation of Translation Made Trivial
by: Zouhar, Vilém, et al.
Published: (2026)
by: Zouhar, Vilém, et al.
Published: (2026)
TAMA: A Human-AI Collaborative Thematic Analysis Framework Using Multi-Agent LLMs for Clinical Interviews
by: Xu, Huimin, et al.
Published: (2025)
by: Xu, Huimin, et al.
Published: (2025)
When Can We Trust LLMs in Mental Health? Large-Scale Benchmarks for Reliable LLM Evaluation
by: Badawi, Abeer, et al.
Published: (2025)
by: Badawi, Abeer, et al.
Published: (2025)
Literary Narrative as Moral Probe : A Cross-System Framework for Evaluating AI Ethical Reasoning and Refusal Behavior
by: Flynn, David C.
Published: (2026)
by: Flynn, David C.
Published: (2026)
A Methodology for Identifying Evaluation Items for Practical Dialogue Systems Based on Business-Dialogue System Alignment Models
by: Nakano, Mikio, et al.
Published: (2026)
by: Nakano, Mikio, et al.
Published: (2026)
ProactiveEval: A Unified Evaluation Framework for Proactive Dialogue Agents
by: Liu, Tianjian, et al.
Published: (2025)
by: Liu, Tianjian, et al.
Published: (2025)
Similar Items
-
Toward Cultural Interpretability: A Linguistic Anthropological Framework for Describing and Evaluating Large Language Models (LLMs)
by: Jones, Graham M., et al.
Published: (2024) -
Pragmatics Meets Culture: Culturally-adapted Artwork Description Generation and Evaluation
by: Zhao, Lingjun, et al.
Published: (2026) -
Questionnaires for Everyone: Streamlining Cross-Cultural Questionnaire Adaptation with GPT-Based Translation Quality Evaluation
by: Haavisto, Otso, et al.
Published: (2024) -
The Case for "Thick Evaluations" of Cultural Representation in AI
by: Qadri, Rida, et al.
Published: (2025) -
DaKultur: Evaluating the Cultural Awareness of Language Models for Danish with Native Speakers
by: Müller-Eberstein, Max, et al.
Published: (2025)