When Tom Eats Kimchi: Evaluating Cultural Bias of Multimodal Large Language Models in Cultural Mixture Contexts
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Kim, Jun Seong, Thu, Kyaw Ye, Ismayilzada, Javad, Park, Junyeong, Kim, Eunsu, Ahmad, Huzama, An, Na Min, Thorne, James, Oh, Alice |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Diffusion Models Through a Global Lens: Are They Culturally Inclusive?
par: Bayramli, Zahra, et autres
Publié: (2025)
par: Bayramli, Zahra, et autres
Publié: (2025)
CLIcK: A Benchmark Dataset of Cultural and Linguistic Intelligence in Korean
par: Kim, Eunsu, et autres
Publié: (2024)
par: Kim, Eunsu, et autres
Publié: (2024)
World in a Frame: Understanding Culture Mixing as a New Challenge for Vision-Language Models
par: Kim, Eunsu, et autres
Publié: (2025)
par: Kim, Eunsu, et autres
Publié: (2025)
Are they lovers or friends? Evaluating LLMs' Social Reasoning in English and Korean Dialogues
par: Kim, Eunsu, et autres
Publié: (2025)
par: Kim, Eunsu, et autres
Publié: (2025)
Flex-TravelPlanner: A Benchmark for Flexible Planning with Language Agents
par: Oh, Juhyun, et autres
Publié: (2025)
par: Oh, Juhyun, et autres
Publié: (2025)
LLM-C3MOD: A Human-LLM Collaborative System for Cross-Cultural Hate Speech Moderation
par: Park, Junyeong, et autres
Publié: (2025)
par: Park, Junyeong, et autres
Publié: (2025)
Multi-FAct: Assessing Factuality of Multilingual LLMs using FActScore
par: Shafayat, Sheikh, et autres
Publié: (2024)
par: Shafayat, Sheikh, et autres
Publié: (2024)
The Generative AI Paradox on Evaluation: What It Can Solve, It May Not Evaluate
par: Oh, Juhyun, et autres
Publié: (2024)
par: Oh, Juhyun, et autres
Publié: (2024)
Designing “Korean” Kimchi: Speculative Configuration of Distance and Commodity Value in the Chinese Kimchi Industry
par: Heangjin Park
Publié: (2025)
par: Heangjin Park
Publié: (2025)
BLUCK: A Benchmark Dataset for Bengali Linguistic Understanding and Cultural Knowledge
par: Kabir, Daeen, et autres
Publié: (2025)
par: Kabir, Daeen, et autres
Publié: (2025)
Physicochemical Property Analyses of Deep‐Frozen Kimchi Cabbage during Long‐Term Storage
par: Dong Hyeon Park, et autres
Publié: (2024)
par: Dong Hyeon Park, et autres
Publié: (2024)
JuICE: A Benchmark for Evaluating LLM-Judge in Identifying Cultural Errors
par: Jin, Jiho, et autres
Publié: (2026)
par: Jin, Jiho, et autres
Publié: (2026)
Spotting Out-of-Character Behavior: Atomic-Level Evaluation of Persona Fidelity in Open-Ended Generation
par: Shin, Jisu, et autres
Publié: (2025)
par: Shin, Jisu, et autres
Publié: (2025)
LLM-as-an-Interviewer: Beyond Static Testing Through Dynamic LLM Evaluation
par: Kim, Eunsu, et autres
Publié: (2024)
par: Kim, Eunsu, et autres
Publié: (2024)
QEVA: A Reference-Free Evaluation Metric for Narrative Video Summarization with Multimodal Question Answering
par: Jung, Woojun, et autres
Publié: (2026)
par: Jung, Woojun, et autres
Publié: (2026)
Survey of Cultural Awareness in Language Models: Text and Beyond
par: Pawar, Siddhesh, et autres
Publié: (2024)
par: Pawar, Siddhesh, et autres
Publié: (2024)
LoCar: Localization-Aware Evaluation of In-Vehicle Assistants through Fine-Grained Sociolinguistic Control
par: Jeong, Seogyeong, et autres
Publié: (2026)
par: Jeong, Seogyeong, et autres
Publié: (2026)
Uncovering Factor Level Preferences to Improve Human-Model Alignment
par: Oh, Juhyun, et autres
Publié: (2024)
par: Oh, Juhyun, et autres
Publié: (2024)
A Dual-Layered Evaluation of Geopolitical and Cultural Bias in LLMs
par: Kim, Sean, et autres
Publié: (2025)
par: Kim, Sean, et autres
Publié: (2025)
Spicy or Not? Exploring Kimchi's Spiciness Perception Across Spicy Food Tolerant and Sensitive Groups
par: Seo‐yeong Chon, et autres
Publié: (2025)
par: Seo‐yeong Chon, et autres
Publié: (2025)
Exposing Blindspots: Cultural Bias Evaluation in Generative Image Models
par: Seo, Huichan, et autres
Publié: (2025)
par: Seo, Huichan, et autres
Publié: (2025)
Culture is Everywhere: A Call for Intentionally Cultural Evaluation
par: Oh, Juhyun, et autres
Publié: (2025)
par: Oh, Juhyun, et autres
Publié: (2025)
MUG-Eval: A Proxy Evaluation Framework for Multilingual Generation Capabilities in Any Language
par: Song, Seyoung, et autres
Publié: (2025)
par: Song, Seyoung, et autres
Publié: (2025)
RoleConflictBench: A Benchmark of Role Conflict Scenarios for Evaluating LLMs' Contextual Sensitivity
par: Shin, Jisu, et autres
Publié: (2025)
par: Shin, Jisu, et autres
Publié: (2025)
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation
par: Kim, Eunsu, et autres
Publié: (2025)
par: Kim, Eunsu, et autres
Publié: (2025)
Understanding EFL Learners' Code-Switching and Teachers' Pedagogical Approaches in LLM-Supported Speaking Practice
par: Park, Junyeong, et autres
Publié: (2025)
par: Park, Junyeong, et autres
Publié: (2025)
AI Should Sense Better, Not Just Scale Bigger: Adaptive Sensing as a Paradigm Shift
par: Baek, Eunsu, et autres
Publié: (2025)
par: Baek, Eunsu, et autres
Publié: (2025)
Unexplored Faces of Robustness and Out-of-Distribution: Covariate Shifts in Environment and Sensor Domains
par: Baek, Eunsu, et autres
Publié: (2024)
par: Baek, Eunsu, et autres
Publié: (2024)
MrSteve: Instruction-Following Agents in Minecraft with What-Where-When Memory
par: Park, Junyeong, et autres
Publié: (2024)
par: Park, Junyeong, et autres
Publié: (2024)
What, When, and Where America Eats
par: Sloan, E
Publié: (2010)
par: Sloan, E
Publié: (2010)
What, When, and Where America Eats
par: Sloan, E. A
Publié: (2012)
par: Sloan, E. A
Publié: (2012)
Context Filtering with Reward Modeling in Question Answering
par: Kim, Sangryul, et autres
Publié: (2024)
par: Kim, Sangryul, et autres
Publié: (2024)
To Eat or Not to Eat
par: Altmann, Peter, et autres
Publié: (2024)
par: Altmann, Peter, et autres
Publié: (2024)
On properness of moduli stacks of $D^{\times}$-shtukas over ramified legs
par: Choi, Yong-Gyu, et autres
Publié: (2025)
par: Choi, Yong-Gyu, et autres
Publié: (2025)
68‐2: A New PWM Micro‐LED Pixel Circuit Using LTPO TFTs with Threshold Voltage and IR‐Drop Compensations
par: Junyeong Kim, et autres
Publié: (2024)
par: Junyeong Kim, et autres
Publié: (2024)
When Will It Fail?: Anomaly to Prompt for Forecasting Future Anomalies in Time Series
par: Park, Min-Yeong, et autres
Publié: (2025)
par: Park, Min-Yeong, et autres
Publié: (2025)
I0T: Embedding Standardization Method Towards Zero Modality Gap
par: An, Na Min, et autres
Publié: (2024)
par: An, Na Min, et autres
Publié: (2024)
Influence of carcass mass on decomposition rate: A medico‐legal entomology perspective
par: Hyeon‐Seok Oh, et autres
Publié: (2024)
par: Hyeon‐Seok Oh, et autres
Publié: (2024)
Enhancing Robustness of Retrieval-Augmented Language Models with In-Context Learning
par: Park, Seong-Il, et autres
Publié: (2024)
par: Park, Seong-Il, et autres
Publié: (2024)
Flavor Lexicon and Testing for Kimchi Available in the United States
par: Jeehyun Lee, et autres
Publié: (2025)
par: Jeehyun Lee, et autres
Publié: (2025)
Documents similaires
-
Diffusion Models Through a Global Lens: Are They Culturally Inclusive?
par: Bayramli, Zahra, et autres
Publié: (2025) -
CLIcK: A Benchmark Dataset of Cultural and Linguistic Intelligence in Korean
par: Kim, Eunsu, et autres
Publié: (2024) -
World in a Frame: Understanding Culture Mixing as a New Challenge for Vision-Language Models
par: Kim, Eunsu, et autres
Publié: (2025) -
Are they lovers or friends? Evaluating LLMs' Social Reasoning in English and Korean Dialogues
par: Kim, Eunsu, et autres
Publié: (2025) -
Flex-TravelPlanner: A Benchmark for Flexible Planning with Language Agents
par: Oh, Juhyun, et autres
Publié: (2025)