Cultural Benchmarking of LLMs in Standard and Dialectal Arabic Dialogues
Fuente:
arXiv
Saved in:
| Main Authors: | Kautsar, Muhammad Dehan Al, Almheiri, Saeed, Ahsan, Momina, Elbouardi, Bilal, Samih, Younes, Ahmad, Sarfraz, Keleg, Amr, Herraoui, Omar El, Elzeky, Kareem, Freihat, Abed Alhakim, Anwar, Mohamed, Xie, Zhuohan, Liang, Junhong, Nasar, Mohammad Rustom Al, Nakov, Preslav, Koto, Fajri |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Instruction-Guided Poetry Generation in Arabic and Its Dialects
by: Sadallah, Abdelrahman, et al.
Published: (2026)
by: Sadallah, Abdelrahman, et al.
Published: (2026)
Linear Semantic Segmentation for Low-Resource Spoken Dialects
by: Chirkunov, Kirill, et al.
Published: (2026)
by: Chirkunov, Kirill, et al.
Published: (2026)
Parallel Tokenizers: Rethinking Vocabulary Design for Cross-Lingual Transfer
by: Kautsar, Muhammad Dehan Al, et al.
Published: (2025)
by: Kautsar, Muhammad Dehan Al, et al.
Published: (2025)
DialectalArabicMMLU: Benchmarking Dialectal Capabilities in Arabic and Multilingual Language Models
by: Altakrori, Malik H., et al.
Published: (2025)
by: Altakrori, Malik H., et al.
Published: (2025)
Toward a Better Localization of Princeton WordNet
by: Freihat, Abed Alhakim
Published: (2025)
by: Freihat, Abed Alhakim
Published: (2025)
What Do Indonesians Really Need from Language Technology? A Nationwide Survey
by: Kautsar, Muhammad Dehan Al, et al.
Published: (2025)
by: Kautsar, Muhammad Dehan Al, et al.
Published: (2025)
Grounding AI-in-Education Development in Teachers' Voices: Findings from a National Survey in Indonesia
by: Aisyah, Nurul, et al.
Published: (2026)
by: Aisyah, Nurul, et al.
Published: (2026)
Role-Aware Language Models for Secure and Contextualized Access Control in Organizations
by: Almheiri, Saeed, et al.
Published: (2025)
by: Almheiri, Saeed, et al.
Published: (2025)
IndoSafety: Culturally Grounded Safety for LLMs in Indonesian Languages
by: Azmi, Muhammad Falensi, et al.
Published: (2025)
by: Azmi, Muhammad Falensi, et al.
Published: (2025)
Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation
by: Hidayat, Naila Shafirni, et al.
Published: (2025)
by: Hidayat, Naila Shafirni, et al.
Published: (2025)
Evaluating Vision-Language and Large Language Models for Automated Student Assessment in Indonesian Classrooms
by: Aisyah, Nurul, et al.
Published: (2025)
by: Aisyah, Nurul, et al.
Published: (2025)
Cross-Cultural Transfer of Commonsense Reasoning in LLMs: Evidence from the Arab World
by: Almheiri, Saeed, et al.
Published: (2025)
by: Almheiri, Saeed, et al.
Published: (2025)
Curriculum Learning and Pseudo-Labeling Improve the Generalization of Multi-Label Arabic Dialect Identification Models
by: Mekky, Ali, et al.
Published: (2026)
by: Mekky, Ali, et al.
Published: (2026)
Advancing the Arabic WordNet: Elevating Content Quality
by: Freihat, Abed Alhakim, et al.
Published: (2024)
by: Freihat, Abed Alhakim, et al.
Published: (2024)
MuDRiC: Multi-Dialect Reasoning for Arabic Commonsense Validation
by: Elozeiri, Kareem, et al.
Published: (2025)
by: Elozeiri, Kareem, et al.
Published: (2025)
Vision Language Models are Confused Tourists
by: Irawan, Patrick Amadeus, et al.
Published: (2025)
by: Irawan, Patrick Amadeus, et al.
Published: (2025)
HALF: Harm-Aware LLM Fairness Evaluation Aligned with Deployment
by: Mekky, Ali, et al.
Published: (2025)
by: Mekky, Ali, et al.
Published: (2025)
Cracking the Code: Multi-domain LLM Evaluation on Real-World Professional Exams in Indonesia
by: Koto, Fajri
Published: (2024)
by: Koto, Fajri
Published: (2024)
Instruction Tuning on Public Government and Cultural Data for Low-Resource Language: a Case Study in Kazakh
by: Laiyk, Nurkhan, et al.
Published: (2025)
by: Laiyk, Nurkhan, et al.
Published: (2025)
SEADialogues: A Multilingual Culturally Grounded Multi-turn Dialogue Dataset on Southeast Asian Languages
by: Kautsar, Muhammad Dehan Al, et al.
Published: (2025)
by: Kautsar, Muhammad Dehan Al, et al.
Published: (2025)
Listen, Correct, and Feed Back: Spoken Pedagogical Feedback Generation
by: Liang, Junhong, et al.
Published: (2026)
by: Liang, Junhong, et al.
Published: (2026)
Preserving Fairness and Safety in Quantized LLMs Through Critical Weight Protection
by: Hakim, Muhammad Alif Al, et al.
Published: (2026)
by: Hakim, Muhammad Alif Al, et al.
Published: (2026)
Controlling Distributional Bias in Multi-Round LLM Generation via KL-Optimized Fine-Tuning
by: Jiang, Yanbei, et al.
Published: (2026)
by: Jiang, Yanbei, et al.
Published: (2026)
Revisiting Common Assumptions about Arabic Dialects in NLP
by: Keleg, Amr, et al.
Published: (2025)
by: Keleg, Amr, et al.
Published: (2025)
AMIR-GRPO: Inducing Implicit Preference Signals into GRPO
by: Yari, Amir Hossein, et al.
Published: (2026)
by: Yari, Amir Hossein, et al.
Published: (2026)
Unveiling Cultural Blind Spots: Analyzing the Limitations of mLLMs in Procedural Text Comprehension
by: Yari, Amir Hossein, et al.
Published: (2025)
by: Yari, Amir Hossein, et al.
Published: (2025)
SAHM: A Benchmark for Arabic Financial and Shari'ah-Compliant Reasoning
by: Elbadry, Rania, et al.
Published: (2026)
by: Elbadry, Rania, et al.
Published: (2026)
Estimating the Level of Dialectness Predicts Interannotator Agreement in Multi-dialect Arabic Datasets
by: Keleg, Amr, et al.
Published: (2024)
by: Keleg, Amr, et al.
Published: (2024)
UrduFactCheck: An Agentic Fact-Checking Framework for Urdu with Evidence Boosting and Benchmarking
by: Ahmad, Sarfraz, et al.
Published: (2025)
by: Ahmad, Sarfraz, et al.
Published: (2025)
FinCARDS: Card-Based Analyst Reranking for Financial Document Question Answering
by: Zhou, Yixi, et al.
Published: (2026)
by: Zhou, Yixi, et al.
Published: (2026)
KazMMLU: Evaluating Language Models on Kazakh, Russian, and Regional Knowledge of Kazakhstan
by: Togmanov, Mukhammed, et al.
Published: (2025)
by: Togmanov, Mukhammed, et al.
Published: (2025)
Enhanced Dark Matter Sensitivity using a Hybrid SiPM-SNSPD-Qubit Detector in Liquid Argon
by: Abed, Faeq, et al.
Published: (2026)
by: Abed, Faeq, et al.
Published: (2026)
Low-Resource Safety Failures Are Action Failures, Not Representation Failures
by: Aziz, Rashad, et al.
Published: (2026)
by: Aziz, Rashad, et al.
Published: (2026)
Multi-Sourced, Multi-Agent Evidence Retrieval for Fact-Checking
by: Gong, Shuzhi, et al.
Published: (2026)
by: Gong, Shuzhi, et al.
Published: (2026)
NeuralNexus at BEA 2025 Shared Task: Retrieval-Augmented Prompting for Mistake Identification in AI Tutors
by: Naeem, Numaan, et al.
Published: (2025)
by: Naeem, Numaan, et al.
Published: (2025)
LLM Alignment for the Arabs: A Homogenous Culture or Diverse Ones?
by: Keleg, Amr
Published: (2025)
by: Keleg, Amr
Published: (2025)
SQLStructEval: Structural Evaluation of LLM Text-to-SQL Generation
by: Zhou, Yixi, et al.
Published: (2026)
by: Zhou, Yixi, et al.
Published: (2026)
The Geometry of Forgetting: Temporal Knowledge Drift as an Independent Axis in LLM Representations
by: Elbadry, Rania, et al.
Published: (2026)
by: Elbadry, Rania, et al.
Published: (2026)
Culturally-Nuanced Story Generation for Reasoning in Low-Resource Languages: The Case of Javanese and Sundanese
by: Pranida, Salsabila Zahirah, et al.
Published: (2025)
by: Pranida, Salsabila Zahirah, et al.
Published: (2025)
IndoCulture: Exploring Geographically-Influenced Cultural Commonsense Reasoning Across Eleven Indonesian Provinces
by: Koto, Fajri, et al.
Published: (2024)
by: Koto, Fajri, et al.
Published: (2024)
Similar Items
-
Instruction-Guided Poetry Generation in Arabic and Its Dialects
by: Sadallah, Abdelrahman, et al.
Published: (2026) -
Linear Semantic Segmentation for Low-Resource Spoken Dialects
by: Chirkunov, Kirill, et al.
Published: (2026) -
Parallel Tokenizers: Rethinking Vocabulary Design for Cross-Lingual Transfer
by: Kautsar, Muhammad Dehan Al, et al.
Published: (2025) -
DialectalArabicMMLU: Benchmarking Dialectal Capabilities in Arabic and Multilingual Language Models
by: Altakrori, Malik H., et al.
Published: (2025) -
Toward a Better Localization of Princeton WordNet
by: Freihat, Abed Alhakim
Published: (2025)