JuICE: A Benchmark for Evaluating LLM-Judge in Identifying Cultural Errors
Fuente:
arXiv
Saved in:
| Main Authors: | Jin, Jiho, Myung, Junho, Oh, Juhyun, Park, Junyeong, Putri, Rifki Afina, Dev, Sunipa, Prabhakaran, Vinodkumar, Oh, Alice |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Can LLM Generate Culturally Relevant Commonsense QA Data? Case Study in Indonesian and Sundanese
by: Putri, Rifki Afina, et al.
Published: (2024)
by: Putri, Rifki Afina, et al.
Published: (2024)
Adapting Language Models to Indonesian Local Languages: An Empirical Study of Language Transferability on Zero-Shot Settings
by: Putri, Rifki Afina
Published: (2025)
by: Putri, Rifki Afina
Published: (2025)
Towards Geo-Culturally Grounded LLM Generations
by: Lertvittayakumjorn, Piyawat, et al.
Published: (2025)
by: Lertvittayakumjorn, Piyawat, et al.
Published: (2025)
Social Bias Benchmark for Generation: A Comparison of Generation and QA-Based Evaluations
by: Jin, Jiho, et al.
Published: (2025)
by: Jin, Jiho, et al.
Published: (2025)
FINEST: Improving LLM Responses to Sensitive Topics Through Fine-Grained Evaluation
by: Oh, Juhyun, et al.
Published: (2026)
by: Oh, Juhyun, et al.
Published: (2026)
SeeGULL Multilingual: a Dataset of Geo-Culturally Situated Stereotypes
by: Bhutani, Mukul, et al.
Published: (2024)
by: Bhutani, Mukul, et al.
Published: (2024)
A Comprehensive Framework to Operationalize Social Stereotypes for Responsible AI Evaluations
by: Davani, Aida, et al.
Published: (2025)
by: Davani, Aida, et al.
Published: (2025)
Cultural Compass: A Framework for Organizing Societal Norms to Detect Violations in Human-AI Conversations
by: Cheng, Myra, et al.
Published: (2026)
by: Cheng, Myra, et al.
Published: (2026)
Survey of Cultural Awareness in Language Models: Text and Beyond
by: Pawar, Siddhesh, et al.
Published: (2024)
by: Pawar, Siddhesh, et al.
Published: (2024)
BEnQA: A Question Answering and Reasoning Benchmark for Bengali and English
by: Shafayat, Sheikh, et al.
Published: (2024)
by: Shafayat, Sheikh, et al.
Published: (2024)
GeniL: A Multilingual Dataset on Generalizing Language
by: Davani, Aida Mostafazadeh, et al.
Published: (2024)
by: Davani, Aida Mostafazadeh, et al.
Published: (2024)
Flex-TravelPlanner: A Benchmark for Flexible Planning with Language Agents
by: Oh, Juhyun, et al.
Published: (2025)
by: Oh, Juhyun, et al.
Published: (2025)
Exploring Cross-Cultural Differences in English Hate Speech Annotations: From Dataset Construction to Analysis
by: Lee, Nayeon, et al.
Published: (2023)
by: Lee, Nayeon, et al.
Published: (2023)
LLM-C3MOD: A Human-LLM Collaborative System for Cross-Cultural Hate Speech Moderation
by: Park, Junyeong, et al.
Published: (2025)
by: Park, Junyeong, et al.
Published: (2025)
Entangled in Representations: Mechanistic Investigation of Cultural Biases in Large Language Models
by: Yu, Haeun, et al.
Published: (2025)
by: Yu, Haeun, et al.
Published: (2025)
SAFARI: A Community-Engaged Approach and Dataset of Stereotype Resources in the Sub-Saharan African Context
by: Verma, Aishwarya, et al.
Published: (2026)
by: Verma, Aishwarya, et al.
Published: (2026)
PapersPlease: A Benchmark for Evaluating Motivational Values of Large Language Models Based on ERG Theory
by: Myung, Junho, et al.
Published: (2025)
by: Myung, Junho, et al.
Published: (2025)
OLA: Output Language Alignment in Code-Switched LLM Interactions
by: Oh, Juhyun, et al.
Published: (2026)
by: Oh, Juhyun, et al.
Published: (2026)
Are they lovers or friends? Evaluating LLMs' Social Reasoning in English and Korean Dialogues
by: Kim, Eunsu, et al.
Published: (2025)
by: Kim, Eunsu, et al.
Published: (2025)
BLEnD: A Benchmark for LLMs on Everyday Knowledge in Diverse Cultures and Languages
by: Myung, Junho, et al.
Published: (2024)
by: Myung, Junho, et al.
Published: (2024)
Culture is Everywhere: A Call for Intentionally Cultural Evaluation
by: Oh, Juhyun, et al.
Published: (2025)
by: Oh, Juhyun, et al.
Published: (2025)
Scaling Cultural Resources for Improving Generative Models
by: Stepanyan, Hayk, et al.
Published: (2025)
by: Stepanyan, Hayk, et al.
Published: (2025)
Multi-FAct: Assessing Factuality of Multilingual LLMs using FActScore
by: Shafayat, Sheikh, et al.
Published: (2024)
by: Shafayat, Sheikh, et al.
Published: (2024)
The Generative AI Paradox on Evaluation: What It Can Solve, It May Not Evaluate
by: Oh, Juhyun, et al.
Published: (2024)
by: Oh, Juhyun, et al.
Published: (2024)
ViSAGe: A Global-Scale Analysis of Visual Stereotypes in Text-to-Image Generation
by: Jha, Akshita, et al.
Published: (2024)
by: Jha, Akshita, et al.
Published: (2024)
Adaptive Data Collection for Latin-American Community-sourced Evaluation of Stereotypes (LACES)
by: Ivetta, Guido, et al.
Published: (2025)
by: Ivetta, Guido, et al.
Published: (2025)
Spotting Out-of-Character Behavior: Atomic-Level Evaluation of Persona Fidelity in Open-Ended Generation
by: Shin, Jisu, et al.
Published: (2025)
by: Shin, Jisu, et al.
Published: (2025)
RoleConflictBench: A Benchmark of Role Conflict Scenarios for Evaluating LLMs' Contextual Sensitivity
by: Shin, Jisu, et al.
Published: (2025)
by: Shin, Jisu, et al.
Published: (2025)
Underpricing of Private Bonds Under Rule 144A*
by: Junho Oh
Published: (2025)
by: Junho Oh
Published: (2025)
When Scaffolding Breaks: Investigating Student Interaction with LLM-Based Writing Support in Real-Time K-12 EFL Classrooms
by: Myung, Junho, et al.
Published: (2025)
by: Myung, Junho, et al.
Published: (2025)
Cultural Authenticity: Comparing LLM Cultural Representations to Native Human Expectations
by: van Liemt, Erin MacMurray, et al.
Published: (2026)
by: van Liemt, Erin MacMurray, et al.
Published: (2026)
Code-Switching In-Context Learning for Cross-Lingual Transfer of Large Language Models
by: Yoo, Haneul, et al.
Published: (2025)
by: Yoo, Haneul, et al.
Published: (2025)
KoBBQ: Korean Bias Benchmark for Question Answering
by: Jin, Jiho, et al.
Published: (2023)
by: Jin, Jiho, et al.
Published: (2023)
Risks of Cultural Erasure in Large Language Models
by: Qadri, Rida, et al.
Published: (2025)
by: Qadri, Rida, et al.
Published: (2025)
Diffusion Models Through a Global Lens: Are They Culturally Inclusive?
by: Bayramli, Zahra, et al.
Published: (2025)
by: Bayramli, Zahra, et al.
Published: (2025)
Domain Generalization for Person Re-identification: A Survey Towards Domain-Agnostic Person Matching
by: Lee, Hyeonseo, et al.
Published: (2025)
by: Lee, Hyeonseo, et al.
Published: (2025)
JuStRank: Benchmarking LLM Judges for System Ranking
by: Gera, Ariel, et al.
Published: (2024)
by: Gera, Ariel, et al.
Published: (2024)
World in a Frame: Understanding Culture Mixing as a New Challenge for Vision-Language Models
by: Kim, Eunsu, et al.
Published: (2025)
by: Kim, Eunsu, et al.
Published: (2025)
D3CODE: Disentangling Disagreements in Data across Cultures on Offensiveness Detection and Evaluation
by: Davani, Aida Mostafazadeh, et al.
Published: (2024)
by: Davani, Aida Mostafazadeh, et al.
Published: (2024)
Taxonomy of User Needs and Actions
by: Shelby, Renee, et al.
Published: (2025)
by: Shelby, Renee, et al.
Published: (2025)
Similar Items
-
Can LLM Generate Culturally Relevant Commonsense QA Data? Case Study in Indonesian and Sundanese
by: Putri, Rifki Afina, et al.
Published: (2024) -
Adapting Language Models to Indonesian Local Languages: An Empirical Study of Language Transferability on Zero-Shot Settings
by: Putri, Rifki Afina
Published: (2025) -
Towards Geo-Culturally Grounded LLM Generations
by: Lertvittayakumjorn, Piyawat, et al.
Published: (2025) -
Social Bias Benchmark for Generation: A Comparison of Generation and QA-Based Evaluations
by: Jin, Jiho, et al.
Published: (2025) -
FINEST: Improving LLM Responses to Sensitive Topics Through Fine-Grained Evaluation
by: Oh, Juhyun, et al.
Published: (2026)