Saved in:
Bibliographic Details
Main Authors: AlOtaibi, Areej, Alyahya, Lina, Alshabanah, Raghad, Alfawzan, Shahad, Alarefei, Shuruq, Alsabti, Reem, Alsubaie, Nouf, Alhuzaymi, Abdulaziz, Alkhelb, Lujain, Alsayari, Majd, Alahmed, Waad, Talabay, Omar, Alowibdi, Jalal, Alelyani, Salem, Bibi, Adel
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2510.13481
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909870643478528
author AlOtaibi, Areej
Alyahya, Lina
Alshabanah, Raghad
Alfawzan, Shahad
Alarefei, Shuruq
Alsabti, Reem
Alsubaie, Nouf
Alhuzaymi, Abdulaziz
Alkhelb, Lujain
Alsayari, Majd
Alahmed, Waad
Talabay, Omar
Alowibdi, Jalal
Alelyani, Salem
Bibi, Adel
author_facet AlOtaibi, Areej
Alyahya, Lina
Alshabanah, Raghad
Alfawzan, Shahad
Alarefei, Shuruq
Alsabti, Reem
Alsubaie, Nouf
Alhuzaymi, Abdulaziz
Alkhelb, Lujain
Alsayari, Majd
Alahmed, Waad
Talabay, Omar
Alowibdi, Jalal
Alelyani, Salem
Bibi, Adel
contents Large Language Models (LLMs) have significantly advanced the field of natural language processing, enhancing capabilities in both language understanding and generation across diverse domains. However, developing LLMs for Arabic presents unique challenges. This paper explores these challenges by focusing on critical aspects such as data curation, tokenizer design, and evaluation. We detail our approach to the collection and filtration of Arabic pre-training datasets, assess the impact of various tokenizer designs on model performance, and examine the limitations of existing Arabic evaluation frameworks, for which we propose a systematic corrective methodology. To promote transparency and facilitate collaborative development, we share our data and methodologies, contributing to the advancement of language modeling, particularly for the Arabic language.
format Preprint
id arxiv_https___arxiv_org_abs_2510_13481
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Tahakom LLM Guidelines and Recipes: From Pre-training Data to an Arabic LLM
AlOtaibi, Areej
Alyahya, Lina
Alshabanah, Raghad
Alfawzan, Shahad
Alarefei, Shuruq
Alsabti, Reem
Alsubaie, Nouf
Alhuzaymi, Abdulaziz
Alkhelb, Lujain
Alsayari, Majd
Alahmed, Waad
Talabay, Omar
Alowibdi, Jalal
Alelyani, Salem
Bibi, Adel
Machine Learning
Large Language Models (LLMs) have significantly advanced the field of natural language processing, enhancing capabilities in both language understanding and generation across diverse domains. However, developing LLMs for Arabic presents unique challenges. This paper explores these challenges by focusing on critical aspects such as data curation, tokenizer design, and evaluation. We detail our approach to the collection and filtration of Arabic pre-training datasets, assess the impact of various tokenizer designs on model performance, and examine the limitations of existing Arabic evaluation frameworks, for which we propose a systematic corrective methodology. To promote transparency and facilitate collaborative development, we share our data and methodologies, contributing to the advancement of language modeling, particularly for the Arabic language.
title Tahakom LLM Guidelines and Recipes: From Pre-training Data to an Arabic LLM
topic Machine Learning
url https://arxiv.org/abs/2510.13481