Cultural Benchmarking of LLMs in Standard and Dialectal Arabic Dialogues

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Kautsar, Muhammad Dehan Al, Almheiri, Saeed, Ahsan, Momina, Elbouardi, Bilal, Samih, Younes, Ahmad, Sarfraz, Keleg, Amr, Herraoui, Omar El, Elzeky, Kareem, Freihat, Abed Alhakim, Anwar, Mohamed, Xie, Zhuohan, Liang, Junhong, Nasar, Mohammad Rustom Al, Nakov, Preslav, Koto, Fajri
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866910183078232064
author Kautsar, Muhammad Dehan Al
Almheiri, Saeed
Ahsan, Momina
Elbouardi, Bilal
Samih, Younes
Ahmad, Sarfraz
Keleg, Amr
Herraoui, Omar El
Elzeky, Kareem
Freihat, Abed Alhakim
Anwar, Mohamed
Xie, Zhuohan
Liang, Junhong
Nasar, Mohammad Rustom Al
Nakov, Preslav
Koto, Fajri
author_facet Kautsar, Muhammad Dehan Al
Almheiri, Saeed
Ahsan, Momina
Elbouardi, Bilal
Samih, Younes
Ahmad, Sarfraz
Keleg, Amr
Herraoui, Omar El
Elzeky, Kareem
Freihat, Abed Alhakim
Anwar, Mohamed
Xie, Zhuohan
Liang, Junhong
Nasar, Mohammad Rustom Al
Nakov, Preslav
Koto, Fajri
contents There is a significant gap in evaluating cultural reasoning in LLMs using conversational datasets that capture culturally rich and dialectal contexts. Most Arabic benchmarks focus on short text snippets in Modern Standard Arabic (MSA), overlooking the cultural nuances that naturally arise in dialogues. To address this gap, we introduce ArabCulture-Dialogue, a culturally grounded conversational dataset covering 13 Arabic-speaking countries, in both MSA and each country's respective dialect, spanning 12 daily-life topics and 54 fine-grained subtopics. We utilize the dataset to form three benchmarking tasks: (i) multiple-choice cultural reasoning, (ii) machine translation between MSA and dialects, and (iii) dialect-steering generation. Our experiments indicate that the performance gap between MSA and Arabic dialects still exists, whereby the models perform worse on all three tasks in the dialectal setup, compared to the MSA one.
format Preprint
id arxiv_https___arxiv_org_abs_2605_00119
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Cultural Benchmarking of LLMs in Standard and Dialectal Arabic Dialogues
Kautsar, Muhammad Dehan Al
Almheiri, Saeed
Ahsan, Momina
Elbouardi, Bilal
Samih, Younes
Ahmad, Sarfraz
Keleg, Amr
Herraoui, Omar El
Elzeky, Kareem
Freihat, Abed Alhakim
Anwar, Mohamed
Xie, Zhuohan
Liang, Junhong
Nasar, Mohammad Rustom Al
Nakov, Preslav
Koto, Fajri
Computation and Language
Artificial Intelligence
I.2.7
There is a significant gap in evaluating cultural reasoning in LLMs using conversational datasets that capture culturally rich and dialectal contexts. Most Arabic benchmarks focus on short text snippets in Modern Standard Arabic (MSA), overlooking the cultural nuances that naturally arise in dialogues. To address this gap, we introduce ArabCulture-Dialogue, a culturally grounded conversational dataset covering 13 Arabic-speaking countries, in both MSA and each country's respective dialect, spanning 12 daily-life topics and 54 fine-grained subtopics. We utilize the dataset to form three benchmarking tasks: (i) multiple-choice cultural reasoning, (ii) machine translation between MSA and dialects, and (iii) dialect-steering generation. Our experiments indicate that the performance gap between MSA and Arabic dialects still exists, whereby the models perform worse on all three tasks in the dialectal setup, compared to the MSA one.
title Cultural Benchmarking of LLMs in Standard and Dialectal Arabic Dialogues
topic Computation and Language
Artificial Intelligence
I.2.7
url https://arxiv.org/abs/2605.00119