ArabicMMLU: Assessing Massive Multitask Language Understanding in Arabic

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Koto, Fajri, Li, Haonan, Shatnawi, Sara, Doughman, Jad, Sadallah, Abdelrahman Boda, Alraeesi, Aisha, Almubarak, Khalid, Alyafeai, Zaid, Sengupta, Neha, Shehata, Shady, Habash, Nizar, Nakov, Preslav, Baldwin, Timothy
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866929441125433344
author Koto, Fajri
Li, Haonan
Shatnawi, Sara
Doughman, Jad
Sadallah, Abdelrahman Boda
Alraeesi, Aisha
Almubarak, Khalid
Alyafeai, Zaid
Sengupta, Neha
Shehata, Shady
Habash, Nizar
Nakov, Preslav
Baldwin, Timothy
author_facet Koto, Fajri
Li, Haonan
Shatnawi, Sara
Doughman, Jad
Sadallah, Abdelrahman Boda
Alraeesi, Aisha
Almubarak, Khalid
Alyafeai, Zaid
Sengupta, Neha
Shehata, Shady
Habash, Nizar
Nakov, Preslav
Baldwin, Timothy
contents The focus of language model evaluation has transitioned towards reasoning and knowledge-intensive tasks, driven by advancements in pretraining large models. While state-of-the-art models are partially trained on large Arabic texts, evaluating their performance in Arabic remains challenging due to the limited availability of relevant datasets. To bridge this gap, we present \datasetname{}, the first multi-task language understanding benchmark for the Arabic language, sourced from school exams across diverse educational levels in different countries spanning North Africa, the Levant, and the Gulf regions. Our data comprises 40 tasks and 14,575 multiple-choice questions in Modern Standard Arabic (MSA) and is carefully constructed by collaborating with native speakers in the region. Our comprehensive evaluations of 35 models reveal substantial room for improvement, particularly among the best open-source models. Notably, BLOOMZ, mT0, LLaMA2, and Falcon struggle to achieve a score of 50%, while even the top-performing Arabic-centric model only achieves a score of 62.3%.
format Preprint
id arxiv_https___arxiv_org_abs_2402_12840
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle ArabicMMLU: Assessing Massive Multitask Language Understanding in Arabic
Koto, Fajri
Li, Haonan
Shatnawi, Sara
Doughman, Jad
Sadallah, Abdelrahman Boda
Alraeesi, Aisha
Almubarak, Khalid
Alyafeai, Zaid
Sengupta, Neha
Shehata, Shady
Habash, Nizar
Nakov, Preslav
Baldwin, Timothy
Computation and Language
The focus of language model evaluation has transitioned towards reasoning and knowledge-intensive tasks, driven by advancements in pretraining large models. While state-of-the-art models are partially trained on large Arabic texts, evaluating their performance in Arabic remains challenging due to the limited availability of relevant datasets. To bridge this gap, we present \datasetname{}, the first multi-task language understanding benchmark for the Arabic language, sourced from school exams across diverse educational levels in different countries spanning North Africa, the Levant, and the Gulf regions. Our data comprises 40 tasks and 14,575 multiple-choice questions in Modern Standard Arabic (MSA) and is carefully constructed by collaborating with native speakers in the region. Our comprehensive evaluations of 35 models reveal substantial room for improvement, particularly among the best open-source models. Notably, BLOOMZ, mT0, LLaMA2, and Falcon struggle to achieve a score of 50%, while even the top-performing Arabic-centric model only achieves a score of 62.3%.
title ArabicMMLU: Assessing Massive Multitask Language Understanding in Arabic
topic Computation and Language
url https://arxiv.org/abs/2402.12840