A Survey of Large Language Models for Arabic Language and its Dialects

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mashaabi, Malak, Al-Khalifa, Shahad, Al-Khalifa, Hend
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913145030705152
author Mashaabi, Malak
Al-Khalifa, Shahad
Al-Khalifa, Hend
author_facet Mashaabi, Malak
Al-Khalifa, Shahad
Al-Khalifa, Hend
contents This survey offers a comprehensive overview of Large Language Models (LLMs) designed for Arabic language and its dialects. It covers key architectures, including encoder-only, decoder-only, and encoder-decoder models, along with the datasets used for pre-training, spanning Classical Arabic, Modern Standard Arabic, and Dialectal Arabic. The study also explores monolingual, bilingual, and multilingual LLMs, analyzing their architectures and performance across downstream tasks, such as sentiment analysis, named entity recognition, and question answering. Furthermore, it assesses the openness of Arabic LLMs based on factors, such as source code availability, training data, model weights, and documentation. The survey highlights the need for more diverse dialectal datasets and attributes the importance of openness for research reproducibility and transparency. It concludes by identifying key challenges and opportunities for future research and stressing the need for more inclusive and representative models.
format Preprint
id arxiv_https___arxiv_org_abs_2410_20238
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle A Survey of Large Language Models for Arabic Language and its Dialects
Mashaabi, Malak
Al-Khalifa, Shahad
Al-Khalifa, Hend
Computation and Language
Artificial Intelligence
This survey offers a comprehensive overview of Large Language Models (LLMs) designed for Arabic language and its dialects. It covers key architectures, including encoder-only, decoder-only, and encoder-decoder models, along with the datasets used for pre-training, spanning Classical Arabic, Modern Standard Arabic, and Dialectal Arabic. The study also explores monolingual, bilingual, and multilingual LLMs, analyzing their architectures and performance across downstream tasks, such as sentiment analysis, named entity recognition, and question answering. Furthermore, it assesses the openness of Arabic LLMs based on factors, such as source code availability, training data, model weights, and documentation. The survey highlights the need for more diverse dialectal datasets and attributes the importance of openness for research reproducibility and transparency. It concludes by identifying key challenges and opportunities for future research and stressing the need for more inclusive and representative models.
title A Survey of Large Language Models for Arabic Language and its Dialects
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2410.20238