DialectalArabicMMLU: Benchmarking Dialectal Capabilities in Arabic and Multilingual Language Models

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Altakrori, Malik H., Habash, Nizar, Freihat, Abed Alhakim, Samih, Younes, Chirkunov, Kirill, AbuOdeh, Muhammed, Florian, Radu, Lynn, Teresa, Nakov, Preslav, Aji, Alham Fikri
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866910062257111040
author Altakrori, Malik H.
Habash, Nizar
Freihat, Abed Alhakim
Samih, Younes
Chirkunov, Kirill
AbuOdeh, Muhammed
Florian, Radu
Lynn, Teresa
Nakov, Preslav
Aji, Alham Fikri
author_facet Altakrori, Malik H.
Habash, Nizar
Freihat, Abed Alhakim
Samih, Younes
Chirkunov, Kirill
AbuOdeh, Muhammed
Florian, Radu
Lynn, Teresa
Nakov, Preslav
Aji, Alham Fikri
contents We present DialectalArabicMMLU, a new benchmark for evaluating the performance of large language models (LLMs) across Arabic dialects. While recently developed Arabic and multilingual benchmarks have advanced LLM evaluation for Modern Standard Arabic (MSA), dialectal varieties remain underrepresented despite their prevalence in everyday communication. DialectalArabicMMLU extends the MMLU-Redux framework through manual translation and adaptation of 3K multiple-choice question-answer pairs into five major dialects (Syrian, Egyptian, Emirati, Saudi, and Moroccan), yielding a total of 15K QA pairs across 32 academic and professional domains (22K QA pairs when also including English and MSA). The benchmark enables systematic assessment of LLM reasoning and comprehension beyond MSA, supporting both task-based and linguistic analysis. We evaluate 19 open-weight Arabic and multilingual LLMs (1B-13B parameters) and report substantial performance variation across dialects, revealing persistent gaps in dialectal generalization. DialectalArabicMMLU provides the first unified, human-curated resource for measuring dialectal understanding in Arabic, thus promoting more inclusive evaluation and future model development.
format Preprint
id arxiv_https___arxiv_org_abs_2510_27543
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DialectalArabicMMLU: Benchmarking Dialectal Capabilities in Arabic and Multilingual Language Models
Altakrori, Malik H.
Habash, Nizar
Freihat, Abed Alhakim
Samih, Younes
Chirkunov, Kirill
AbuOdeh, Muhammed
Florian, Radu
Lynn, Teresa
Nakov, Preslav
Aji, Alham Fikri
Computation and Language
Artificial Intelligence
We present DialectalArabicMMLU, a new benchmark for evaluating the performance of large language models (LLMs) across Arabic dialects. While recently developed Arabic and multilingual benchmarks have advanced LLM evaluation for Modern Standard Arabic (MSA), dialectal varieties remain underrepresented despite their prevalence in everyday communication. DialectalArabicMMLU extends the MMLU-Redux framework through manual translation and adaptation of 3K multiple-choice question-answer pairs into five major dialects (Syrian, Egyptian, Emirati, Saudi, and Moroccan), yielding a total of 15K QA pairs across 32 academic and professional domains (22K QA pairs when also including English and MSA). The benchmark enables systematic assessment of LLM reasoning and comprehension beyond MSA, supporting both task-based and linguistic analysis. We evaluate 19 open-weight Arabic and multilingual LLMs (1B-13B parameters) and report substantial performance variation across dialects, revealing persistent gaps in dialectal generalization. DialectalArabicMMLU provides the first unified, human-curated resource for measuring dialectal understanding in Arabic, thus promoting more inclusive evaluation and future model development.
title DialectalArabicMMLU: Benchmarking Dialectal Capabilities in Arabic and Multilingual Language Models
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2510.27543