MAKIEval: A Multilingual Automatic WiKidata-based Framework for Cultural Awareness Evaluation for LLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhao, Raoyuan, Chen, Beiduo, Plank, Barbara, Hedderich, Michael A. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A Comprehensive Evaluation of Multilingual Chain-of-Thought Reasoning: Performance, Consistency, and Faithfulness Across Languages
von: Zhao, Raoyuan, et al.
Veröffentlicht: (2025)
von: Zhao, Raoyuan, et al.
Veröffentlicht: (2025)
Crosslingual On-Policy Self-Distillation for Multilingual Reasoning
von: Liu, Yihong, et al.
Veröffentlicht: (2026)
von: Liu, Yihong, et al.
Veröffentlicht: (2026)
Large Reasoning Models Are (Not Yet) Multilingual Latent Reasoners
von: Liu, Yihong, et al.
Veröffentlicht: (2026)
von: Liu, Yihong, et al.
Veröffentlicht: (2026)
Evaluating Robustness of Large Language Models Against Multilingual Typographical Errors
von: Zhao, Raoyuan, et al.
Veröffentlicht: (2025)
von: Zhao, Raoyuan, et al.
Veröffentlicht: (2025)
Probing LLMs for Multilingual Discourse Generalization Through a Unified Label Set
von: Eichin, Florian, et al.
Veröffentlicht: (2025)
von: Eichin, Florian, et al.
Veröffentlicht: (2025)
What's the Difference? Supporting Users in Identifying the Effects of Prompt and Model Changes Through Token Patterns
von: Hedderich, Michael A., et al.
Veröffentlicht: (2025)
von: Hedderich, Michael A., et al.
Veröffentlicht: (2025)
Do We Know What LLMs Don't Know? A Study of Consistency in Knowledge Probing
von: Zhao, Raoyuan, et al.
Veröffentlicht: (2025)
von: Zhao, Raoyuan, et al.
Veröffentlicht: (2025)
A Rose by Any Other Name: LLM-Generated Explanations Are Good Proxies for Human Explanations to Collect Label Distributions on NLI
von: Chen, Beiduo, et al.
Veröffentlicht: (2024)
von: Chen, Beiduo, et al.
Veröffentlicht: (2024)
Disagreeing Rationales: Rethinking Classification and Explainability Evaluation in Hate Speech Detection
von: Muscato, Benedetta, et al.
Veröffentlicht: (2026)
von: Muscato, Benedetta, et al.
Veröffentlicht: (2026)
Reasoning that Travels: Dissecting How Chain-of-Thought Transfers Across Models
von: Cheng, Xinyuan, et al.
Veröffentlicht: (2026)
von: Cheng, Xinyuan, et al.
Veröffentlicht: (2026)
Copy First, Translate Later: Interpreting Translation Dynamics in Multilingual Pretraining
von: Körner, Felicia, et al.
Veröffentlicht: (2026)
von: Körner, Felicia, et al.
Veröffentlicht: (2026)
ReverseMath: Answer Inversion for Scalable and Verifiable Mathematical Problem Generation
von: Zhao, Raoyuan, et al.
Veröffentlicht: (2026)
von: Zhao, Raoyuan, et al.
Veröffentlicht: (2026)
Beyond Input Understanding: Diagnosing Multilingual Mathematical Reasoning with Directed Acyclic Trace Graphs
von: Zhang, Jiaqiao, et al.
Veröffentlicht: (2026)
von: Zhang, Jiaqiao, et al.
Veröffentlicht: (2026)
Threading the Needle: Reweaving Chain-of-Thought Reasoning to Explain Human Label Variation
von: Chen, Beiduo, et al.
Veröffentlicht: (2025)
von: Chen, Beiduo, et al.
Veröffentlicht: (2025)
"Seeing the Big through the Small": Can LLMs Approximate Human Judgment Distributions on NLI from a Few Explanations?
von: Chen, Beiduo, et al.
Veröffentlicht: (2024)
von: Chen, Beiduo, et al.
Veröffentlicht: (2024)
Understanding When Tree of Thoughts Succeeds: Larger Models Excel in Generation, Not Discrimination
von: Chen, Qiqi, et al.
Veröffentlicht: (2024)
von: Chen, Qiqi, et al.
Veröffentlicht: (2024)
Decoupling the Effect of Chain-of-Thought Reasoning: A Human Label Variation Perspective
von: Chen, Beiduo, et al.
Veröffentlicht: (2026)
von: Chen, Beiduo, et al.
Veröffentlicht: (2026)
LiTEx: A Linguistic Taxonomy of Explanations for Understanding Within-Label Variation in Natural Language Inference
von: Hong, Pingjun, et al.
Veröffentlicht: (2025)
von: Hong, Pingjun, et al.
Veröffentlicht: (2025)
Human Label Variation as Stable Signal: Learning Annotator-Specific Explanation Behavior via Cross-Annotator Preference Optimization
von: Chen, Beiduo, et al.
Veröffentlicht: (2026)
von: Chen, Beiduo, et al.
Veröffentlicht: (2026)
The Potential and Challenges of Evaluating Attitudes, Opinions, and Values in Large Language Models
von: Ma, Bolei, et al.
Veröffentlicht: (2024)
von: Ma, Bolei, et al.
Veröffentlicht: (2024)
MCEval: A Dynamic Framework for Fair Multilingual Cultural Evaluation of LLMs
von: Huang, Shulin, et al.
Veröffentlicht: (2025)
von: Huang, Shulin, et al.
Veröffentlicht: (2025)
Agree, Disagree, Explain: Decomposing Human Label Variation in NLI through the Lens of Explanations
von: Hong, Pingjun, et al.
Veröffentlicht: (2025)
von: Hong, Pingjun, et al.
Veröffentlicht: (2025)
Human Uncertainty-Aware Data Selection and Automatic Labeling in Visual Question Answering
von: Lan, Jian, et al.
Veröffentlicht: (2025)
von: Lan, Jian, et al.
Veröffentlicht: (2025)
Multilingual != Multicultural: Evaluating Gaps Between Multilingual Capabilities and Cultural Alignment in LLMs
von: Rystrøm, Jonathan, et al.
Veröffentlicht: (2025)
von: Rystrøm, Jonathan, et al.
Veröffentlicht: (2025)
Do LLMs Give Psychometrically Plausible Responses in Educational Assessments?
von: Säuberli, Andreas, et al.
Veröffentlicht: (2025)
von: Säuberli, Andreas, et al.
Veröffentlicht: (2025)
Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey
von: Mondorf, Philipp, et al.
Veröffentlicht: (2024)
von: Mondorf, Philipp, et al.
Veröffentlicht: (2024)
Culturally-Aware Conversations: A Framework & Benchmark for LLMs
von: Havaldar, Shreya, et al.
Veröffentlicht: (2025)
von: Havaldar, Shreya, et al.
Veröffentlicht: (2025)
RAcQUEt: Unveiling the Dangers of Overlooked Referential Ambiguity in Visual LLMs
von: Testoni, Alberto, et al.
Veröffentlicht: (2024)
von: Testoni, Alberto, et al.
Veröffentlicht: (2024)
BoN Appetit Team at LeWiDi-2025: Best-of-N Test-time Scaling Can Not Stomach Annotation Disagreements (Yet)
von: Ruiz, Tomas, et al.
Veröffentlicht: (2025)
von: Ruiz, Tomas, et al.
Veröffentlicht: (2025)
CLIMATELI: Evaluating Entity Linking on Climate Change Data
von: Zhou, Shijia, et al.
Veröffentlicht: (2024)
von: Zhou, Shijia, et al.
Veröffentlicht: (2024)
When Meanings Meet: Investigating the Emergence and Quality of Shared Concept Spaces during Multilingual Language Model Training
von: Körner, Felicia, et al.
Veröffentlicht: (2026)
von: Körner, Felicia, et al.
Veröffentlicht: (2026)
The Call for Socially Aware Language Technologies
von: Yang, Diyi, et al.
Veröffentlicht: (2024)
von: Yang, Diyi, et al.
Veröffentlicht: (2024)
Dialetto, ma Quanto Dialetto? Transcribing and Evaluating Dialects on a Continuum
von: Shim, Ryan Soh-Eun, et al.
Veröffentlicht: (2024)
von: Shim, Ryan Soh-Eun, et al.
Veröffentlicht: (2024)
MEDAL: A Framework for Benchmarking LLMs as Multilingual Open-Domain Dialogue Evaluators
von: Mendonça, John, et al.
Veröffentlicht: (2025)
von: Mendonça, John, et al.
Veröffentlicht: (2025)
Are Multilingual LLMs Culturally-Diverse Reasoners? An Investigation into Multicultural Proverbs and Sayings
von: Liu, Chen Cecilia, et al.
Veröffentlicht: (2023)
von: Liu, Chen Cecilia, et al.
Veröffentlicht: (2023)
MMA-ASIA: A Multilingual and Multimodal Alignment Framework for Culturally-Grounded Evaluation
von: Zheng, Weihua, et al.
Veröffentlicht: (2025)
von: Zheng, Weihua, et al.
Veröffentlicht: (2025)
Liar, Liar, Logical Mire: A Benchmark for Suppositional Reasoning in Large Language Models
von: Mondorf, Philipp, et al.
Veröffentlicht: (2024)
von: Mondorf, Philipp, et al.
Veröffentlicht: (2024)
Disentangling Language and Culture for Evaluating Multilingual Large Language Models
von: Ying, Jiahao, et al.
Veröffentlicht: (2025)
von: Ying, Jiahao, et al.
Veröffentlicht: (2025)
CARE: Multilingual Human Preference Learning for Cultural Awareness
von: Guo, Geyang, et al.
Veröffentlicht: (2025)
von: Guo, Geyang, et al.
Veröffentlicht: (2025)
CURE: Cultural Understanding and Reasoning Evaluation - A Framework for "Thick" Culture Alignment Evaluation in LLMs
von: Vo, Truong, et al.
Veröffentlicht: (2025)
von: Vo, Truong, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
A Comprehensive Evaluation of Multilingual Chain-of-Thought Reasoning: Performance, Consistency, and Faithfulness Across Languages
von: Zhao, Raoyuan, et al.
Veröffentlicht: (2025) -
Crosslingual On-Policy Self-Distillation for Multilingual Reasoning
von: Liu, Yihong, et al.
Veröffentlicht: (2026) -
Large Reasoning Models Are (Not Yet) Multilingual Latent Reasoners
von: Liu, Yihong, et al.
Veröffentlicht: (2026) -
Evaluating Robustness of Large Language Models Against Multilingual Typographical Errors
von: Zhao, Raoyuan, et al.
Veröffentlicht: (2025) -
Probing LLMs for Multilingual Discourse Generalization Through a Unified Label Set
von: Eichin, Florian, et al.
Veröffentlicht: (2025)