A Structured Framework for Evaluating and Enhancing Interpretive Capabilities of Multimodal LLMs in Culturally Situated Tasks
Fuente:
arXiv
Saved in:
| Main Authors: | Yu, Haorui, Ruiz-Dolz, Ramon, Yi, Qiufeng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Cross-Cultural Expert-Level Art Critique Evaluation with Vision-Language Models
by: Yu, Haorui, et al.
Published: (2026)
by: Yu, Haorui, et al.
Published: (2026)
VULCA-Bench: A Multicultural Vision-Language Benchmark for Evaluating Cultural Understanding
by: Yu, Haorui, et al.
Published: (2026)
by: Yu, Haorui, et al.
Published: (2026)
Seeing Symbols, Missing Cultures: Probing Vision-Language Models' Reasoning on Fire Imagery and Cultural Meaning
by: Yu, Haorui, et al.
Published: (2025)
by: Yu, Haorui, et al.
Published: (2025)
An Explainable Framework for Misinformation Identification via Critical Question Answering
by: Ruiz-Dolz, Ramon, et al.
Published: (2025)
by: Ruiz-Dolz, Ramon, et al.
Published: (2025)
Automatic Debate Evaluation with Argumentation Semantics and Natural Language Argument Graph Networks
by: Ruiz-Dolz, Ramon, et al.
Published: (2022)
by: Ruiz-Dolz, Ramon, et al.
Published: (2022)
VivesDebate-Speech: A Corpus of Spoken Argumentation to Leverage Audio Features for Argument Mining
by: Ruiz-Dolz, Ramon, et al.
Published: (2023)
by: Ruiz-Dolz, Ramon, et al.
Published: (2023)
Hidden in Plain Sight: Evaluation of the Deception Detection Capabilities of LLMs in Multimodal Settings
by: Miah, Md Messal Monem, et al.
Published: (2025)
by: Miah, Md Messal Monem, et al.
Published: (2025)
Evaluating Linguistic Capabilities of Multimodal LLMs in the Lens of Few-Shot Learning
by: Dogan, Mustafa, et al.
Published: (2024)
by: Dogan, Mustafa, et al.
Published: (2024)
NLAS-multi: A Multilingual Corpus of Automatically Generated Natural Language Argumentation Schemes
by: Ruiz-Dolz, Ramon, et al.
Published: (2024)
by: Ruiz-Dolz, Ramon, et al.
Published: (2024)
Multilingual != Multicultural: Evaluating Gaps Between Multilingual Capabilities and Cultural Alignment in LLMs
by: Rystrøm, Jonathan, et al.
Published: (2025)
by: Rystrøm, Jonathan, et al.
Published: (2025)
Toward Cultural Interpretability: A Linguistic Anthropological Framework for Describing and Evaluating Large Language Models (LLMs)
by: Jones, Graham M., et al.
Published: (2024)
by: Jones, Graham M., et al.
Published: (2024)
Multimodal Cultural Safety: Evaluation Framework and Alignment Strategies
by: Qiu, Haoyi, et al.
Published: (2025)
by: Qiu, Haoyi, et al.
Published: (2025)
CURE: Cultural Understanding and Reasoning Evaluation - A Framework for "Thick" Culture Alignment Evaluation in LLMs
by: Vo, Truong, et al.
Published: (2025)
by: Vo, Truong, et al.
Published: (2025)
MCEval: A Dynamic Framework for Fair Multilingual Cultural Evaluation of LLMs
by: Huang, Shulin, et al.
Published: (2025)
by: Huang, Shulin, et al.
Published: (2025)
ShredBench: Evaluating the Semantic Reasoning Capabilities of Multimodal LLMs in Document Reconstruction
by: Guo, Zichun, et al.
Published: (2026)
by: Guo, Zichun, et al.
Published: (2026)
Understanding Multimodal LLMs: the Mechanistic Interpretability of Llava in Visual Question Answering
by: Yu, Zeping, et al.
Published: (2024)
by: Yu, Zeping, et al.
Published: (2024)
XCR-Bench: A Multi-Task Benchmark for Evaluating Cultural Reasoning in LLMs
by: Kabir, Mohsinul, et al.
Published: (2026)
by: Kabir, Mohsinul, et al.
Published: (2026)
Dual Perspectives in Emotion Attribution: A Generator-Interpreter Framework for Cross-Cultural Analysis of Emotion in LLMs
by: Turdubaeva, Aizirek, et al.
Published: (2026)
by: Turdubaeva, Aizirek, et al.
Published: (2026)
A Framework for Evaluating LLMs Under Task Indeterminacy
by: Guerdan, Luke, et al.
Published: (2024)
by: Guerdan, Luke, et al.
Published: (2024)
SCAN: Structured Capability Assessment and Navigation for LLMs
by: Wang, Zongqi, et al.
Published: (2025)
by: Wang, Zongqi, et al.
Published: (2025)
Time Series Forecasting with LLMs: Understanding and Enhancing Model Capabilities
by: Tang, Hua, et al.
Published: (2024)
by: Tang, Hua, et al.
Published: (2024)
Multimodal Situational Safety
by: Zhou, Kaiwen, et al.
Published: (2024)
by: Zhou, Kaiwen, et al.
Published: (2024)
LLMs and their Limited Theory of Mind: Evaluating Mental State Annotations in Situated Dialogue
by: Kowalyshyn, Katharine, et al.
Published: (2025)
by: Kowalyshyn, Katharine, et al.
Published: (2025)
FOFO: A Benchmark to Evaluate LLMs' Format-Following Capability
by: Xia, Congying, et al.
Published: (2024)
by: Xia, Congying, et al.
Published: (2024)
Code-Vision: Evaluating Multimodal LLMs Logic Understanding and Code Generation Capabilities
by: Wang, Hanbin, et al.
Published: (2025)
by: Wang, Hanbin, et al.
Published: (2025)
MMA-ASIA: A Multilingual and Multimodal Alignment Framework for Culturally-Grounded Evaluation
by: Zheng, Weihua, et al.
Published: (2025)
by: Zheng, Weihua, et al.
Published: (2025)
FEEL: A Framework for Evaluating Emotional Support Capability with Large Language Models
by: Zhang, Huaiwen, et al.
Published: (2024)
by: Zhang, Huaiwen, et al.
Published: (2024)
CDT: A Comprehensive Capability Framework for Large Language Models Across Cognition, Domain, and Task
by: Mo, Haosi, et al.
Published: (2025)
by: Mo, Haosi, et al.
Published: (2025)
"A Woman is More Culturally Knowledgeable than A Man?": The Effect of Personas on Cultural Norm Interpretation in LLMs
by: Kamruzzaman, Mahammed, et al.
Published: (2024)
by: Kamruzzaman, Mahammed, et al.
Published: (2024)
HIPPO: Enhancing the Table Understanding Capability of LLMs through Hybrid-Modal Preference Optimization
by: Wang, Haolan, et al.
Published: (2025)
by: Wang, Haolan, et al.
Published: (2025)
Enhancing Function-Calling Capabilities in LLMs: Strategies for Prompt Formats, Data Integration, and Multilingual Translation
by: Chen, Yi-Chang, et al.
Published: (2024)
by: Chen, Yi-Chang, et al.
Published: (2024)
MAKIEval: A Multilingual Automatic WiKidata-based Framework for Cultural Awareness Evaluation for LLMs
by: Zhao, Raoyuan, et al.
Published: (2025)
by: Zhao, Raoyuan, et al.
Published: (2025)
Multimodal Task Interference: A Benchmark and Analysis of History-Target Mismatch in Multimodal LLMs
by: Kawarada, Masayuki, et al.
Published: (2026)
by: Kawarada, Masayuki, et al.
Published: (2026)
Soohak: A Mathematician-Curated Benchmark for Evaluating Research-level Math Capabilities of LLMs
by: Son, Guijin, et al.
Published: (2026)
by: Son, Guijin, et al.
Published: (2026)
Culturally-Aware Conversations: A Framework & Benchmark for LLMs
by: Havaldar, Shreya, et al.
Published: (2025)
by: Havaldar, Shreya, et al.
Published: (2025)
Grounding Multilingual Multimodal LLMs With Cultural Knowledge
by: Nyandwi, Jean de Dieu, et al.
Published: (2025)
by: Nyandwi, Jean de Dieu, et al.
Published: (2025)
Defining Cultural Capabilities for AI Evaluation: A Taxonomy Grounded in Intercultural Communication Theory
by: Nejadgholi, Isar, et al.
Published: (2026)
by: Nejadgholi, Isar, et al.
Published: (2026)
SO-Bench: A Structural Output Evaluation of Multimodal LLMs
by: Feng, Di, et al.
Published: (2025)
by: Feng, Di, et al.
Published: (2025)
STEM: Efficient Relative Capability Evaluation of LLMs through Structured Transition Samples
by: Hu, Haiquan, et al.
Published: (2025)
by: Hu, Haiquan, et al.
Published: (2025)
Interpretability Framework for LLMs in Undergraduate Calculus
by: Dakshit, Sagnik, et al.
Published: (2025)
by: Dakshit, Sagnik, et al.
Published: (2025)
Similar Items
-
Cross-Cultural Expert-Level Art Critique Evaluation with Vision-Language Models
by: Yu, Haorui, et al.
Published: (2026) -
VULCA-Bench: A Multicultural Vision-Language Benchmark for Evaluating Cultural Understanding
by: Yu, Haorui, et al.
Published: (2026) -
Seeing Symbols, Missing Cultures: Probing Vision-Language Models' Reasoning on Fire Imagery and Cultural Meaning
by: Yu, Haorui, et al.
Published: (2025) -
An Explainable Framework for Misinformation Identification via Critical Question Answering
by: Ruiz-Dolz, Ramon, et al.
Published: (2025) -
Automatic Debate Evaluation with Argumentation Semantics and Natural Language Argument Graph Networks
by: Ruiz-Dolz, Ramon, et al.
Published: (2022)