Unveiling Cultural Blind Spots: Analyzing the Limitations of mLLMs in Procedural Text Comprehension
Fuente:
arXiv
Saved in:
| Main Authors: | Yari, Amir Hossein, Koto, Fajri |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Revisiting Metric Reliability for Fine-grained Evaluation of Machine Translation and Summarization in Indian Languages
by: Yari, Amir Hossein, et al.
Published: (2025)
by: Yari, Amir Hossein, et al.
Published: (2025)
AMIR-GRPO: Inducing Implicit Preference Signals into GRPO
by: Yari, Amir Hossein, et al.
Published: (2026)
by: Yari, Amir Hossein, et al.
Published: (2026)
Cracking the Code: Multi-domain LLM Evaluation on Real-World Professional Exams in Indonesia
by: Koto, Fajri
Published: (2024)
by: Koto, Fajri
Published: (2024)
LLMs as Cultural Archives: Cultural Commonsense Knowledge Graph Extraction
by: Tonga, Junior Cedric, et al.
Published: (2026)
by: Tonga, Junior Cedric, et al.
Published: (2026)
Are Multilingual LLMs Culturally-Diverse Reasoners? An Investigation into Multicultural Proverbs and Sayings
by: Liu, Chen Cecilia, et al.
Published: (2023)
by: Liu, Chen Cecilia, et al.
Published: (2023)
IndoCulture: Exploring Geographically-Influenced Cultural Commonsense Reasoning Across Eleven Indonesian Provinces
by: Koto, Fajri, et al.
Published: (2024)
by: Koto, Fajri, et al.
Published: (2024)
IndoSafety: Culturally Grounded Safety for LLMs in Indonesian Languages
by: Azmi, Muhammad Falensi, et al.
Published: (2025)
by: Azmi, Muhammad Falensi, et al.
Published: (2025)
Preserving Fairness and Safety in Quantized LLMs Through Critical Weight Protection
by: Hakim, Muhammad Alif Al, et al.
Published: (2026)
by: Hakim, Muhammad Alif Al, et al.
Published: (2026)
Parallel Tokenizers: Rethinking Vocabulary Design for Cross-Lingual Transfer
by: Kautsar, Muhammad Dehan Al, et al.
Published: (2025)
by: Kautsar, Muhammad Dehan Al, et al.
Published: (2025)
Cross-Cultural Transfer of Commonsense Reasoning in LLMs: Evidence from the Arab World
by: Almheiri, Saeed, et al.
Published: (2025)
by: Almheiri, Saeed, et al.
Published: (2025)
Culturally-Nuanced Story Generation for Reasoning in Low-Resource Languages: The Case of Javanese and Sundanese
by: Pranida, Salsabila Zahirah, et al.
Published: (2025)
by: Pranida, Salsabila Zahirah, et al.
Published: (2025)
IndoBias: A Dual Track Culturally Grounded Benchmark for LLMs Bias Evaluation in Indonesian Languages
by: Hanif, Ikhlasul Akmal, et al.
Published: (2026)
by: Hanif, Ikhlasul Akmal, et al.
Published: (2026)
ZPD-SCA: Unveiling the Blind Spots of LLMs in Assessing Students' Cognitive Abilities
by: Dong, Wenhan, et al.
Published: (2025)
by: Dong, Wenhan, et al.
Published: (2025)
Low-Resource Safety Failures Are Action Failures, Not Representation Failures
by: Aziz, Rashad, et al.
Published: (2026)
by: Aziz, Rashad, et al.
Published: (2026)
Exploring Language-Agnosticity in Function Vectors: A Case Study in Machine Translation
by: Laiyk, Nurkhan, et al.
Published: (2026)
by: Laiyk, Nurkhan, et al.
Published: (2026)
Listen, Correct, and Feed Back: Spoken Pedagogical Feedback Generation
by: Liang, Junhong, et al.
Published: (2026)
by: Liang, Junhong, et al.
Published: (2026)
Instruction Tuning on Public Government and Cultural Data for Low-Resource Language: a Case Study in Kazakh
by: Laiyk, Nurkhan, et al.
Published: (2025)
by: Laiyk, Nurkhan, et al.
Published: (2025)
Finding Blind Spots in Evaluator LLMs with Interpretable Checklists
by: Doddapaneni, Sumanth, et al.
Published: (2024)
by: Doddapaneni, Sumanth, et al.
Published: (2024)
What Do Indonesians Really Need from Language Technology? A Nationwide Survey
by: Kautsar, Muhammad Dehan Al, et al.
Published: (2025)
by: Kautsar, Muhammad Dehan Al, et al.
Published: (2025)
Grounding AI-in-Education Development in Teachers' Voices: Findings from a National Survey in Indonesia
by: Aisyah, Nurul, et al.
Published: (2026)
by: Aisyah, Nurul, et al.
Published: (2026)
Zero-shot Sentiment Analysis in Low-Resource Languages Using a Multilingual Sentiment Lexicon
by: Koto, Fajri, et al.
Published: (2024)
by: Koto, Fajri, et al.
Published: (2024)
Simple Linguistic Inferences of Large Language Models (LLMs): Blind Spots and Blinds
by: Basmov, Victoria, et al.
Published: (2023)
by: Basmov, Victoria, et al.
Published: (2023)
Commonsense Reasoning in Arab Culture
by: Sadallah, Abdelrahman, et al.
Published: (2025)
by: Sadallah, Abdelrahman, et al.
Published: (2025)
The Rarity Blind Spot: A Framework for Evaluating Statistical Reasoning in LLMs
by: Maekawa, Seiji, et al.
Published: (2025)
by: Maekawa, Seiji, et al.
Published: (2025)
Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation
by: Hidayat, Naila Shafirni, et al.
Published: (2025)
by: Hidayat, Naila Shafirni, et al.
Published: (2025)
Evaluating Vision-Language and Large Language Models for Automated Student Assessment in Indonesian Classrooms
by: Aisyah, Nurul, et al.
Published: (2025)
by: Aisyah, Nurul, et al.
Published: (2025)
Temporal Blind Spots in Large Language Models
by: Wallat, Jonas, et al.
Published: (2024)
by: Wallat, Jonas, et al.
Published: (2024)
Simulating LLM-to-LLM Tutoring for Multilingual Math Feedback
by: Tonga, Junior Cedric, et al.
Published: (2025)
by: Tonga, Junior Cedric, et al.
Published: (2025)
Controlling Distributional Bias in Multi-Round LLM Generation via KL-Optimized Fine-Tuning
by: Jiang, Yanbei, et al.
Published: (2026)
by: Jiang, Yanbei, et al.
Published: (2026)
CMMLU: Measuring massive multitask language understanding in Chinese
by: Li, Haonan, et al.
Published: (2023)
by: Li, Haonan, et al.
Published: (2023)
Syntactic Blind Spots: How Misalignment Leads to LLMs Mathematical Errors
by: Williamson, Dane, et al.
Published: (2025)
by: Williamson, Dane, et al.
Published: (2025)
Linguistic Blind Spots in Clinical Decision Extraction
by: Elgaar, Mohamed, et al.
Published: (2026)
by: Elgaar, Mohamed, et al.
Published: (2026)
Cultural Benchmarking of LLMs in Standard and Dialectal Arabic Dialogues
by: Kautsar, Muhammad Dehan Al, et al.
Published: (2026)
by: Kautsar, Muhammad Dehan Al, et al.
Published: (2026)
Macaron: Controlled, Human-Written Benchmark for Multilingual and Multicultural Reasoning via Template-Filling
by: Elsetohy, Alaa, et al.
Published: (2026)
by: Elsetohy, Alaa, et al.
Published: (2026)
Fluent but Unfeeling: The Emotional Blind Spots of Language Models
by: Shu, Bangzhao, et al.
Published: (2025)
by: Shu, Bangzhao, et al.
Published: (2025)
SEADialogues: A Multilingual Culturally Grounded Multi-turn Dialogue Dataset on Southeast Asian Languages
by: Kautsar, Muhammad Dehan Al, et al.
Published: (2025)
by: Kautsar, Muhammad Dehan Al, et al.
Published: (2025)
From Text to Emotion: Unveiling the Emotion Annotation Capabilities of LLMs
by: Niu, Minxue, et al.
Published: (2024)
by: Niu, Minxue, et al.
Published: (2024)
Positional Failures in Long-Context LLMs: A Blind Spot in Reasoning Benchmarks
by: Zhang, Chuyifei, et al.
Published: (2026)
by: Zhang, Chuyifei, et al.
Published: (2026)
Linguistic Blind Spots of Large Language Models
by: Cheng, Jiali, et al.
Published: (2025)
by: Cheng, Jiali, et al.
Published: (2025)
CoBia: Constructed Conversations Can Trigger Otherwise Concealed Societal Biases in LLMs
by: Nikeghbal, Nafiseh, et al.
Published: (2025)
by: Nikeghbal, Nafiseh, et al.
Published: (2025)
Similar Items
-
Revisiting Metric Reliability for Fine-grained Evaluation of Machine Translation and Summarization in Indian Languages
by: Yari, Amir Hossein, et al.
Published: (2025) -
AMIR-GRPO: Inducing Implicit Preference Signals into GRPO
by: Yari, Amir Hossein, et al.
Published: (2026) -
Cracking the Code: Multi-domain LLM Evaluation on Real-World Professional Exams in Indonesia
by: Koto, Fajri
Published: (2024) -
LLMs as Cultural Archives: Cultural Commonsense Knowledge Graph Extraction
by: Tonga, Junior Cedric, et al.
Published: (2026) -
Are Multilingual LLMs Culturally-Diverse Reasoners? An Investigation into Multicultural Proverbs and Sayings
by: Liu, Chen Cecilia, et al.
Published: (2023)