Kuwain 1.5B: An Arabic SLM via Language Injection

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Hennara, Khalil, Chrouf, Sara, Hamed, Mohamed Motaism, Aldallal, Zeina, Hadid, Omar, AlModhayan, Safwan
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866915453752836096
author Hennara, Khalil
Chrouf, Sara
Hamed, Mohamed Motaism
Aldallal, Zeina
Hadid, Omar
AlModhayan, Safwan
author_facet Hennara, Khalil
Chrouf, Sara
Hamed, Mohamed Motaism
Aldallal, Zeina
Hadid, Omar
AlModhayan, Safwan
contents Enhancing existing models with new knowledge is a crucial aspect of AI development. This paper introduces a novel method for integrating a new language into a large language model (LLM). Our approach successfully incorporates a previously unseen target language into an existing LLM without compromising its prior knowledge. We trained a tiny model with 1.5 billion parameters named Kuwain by injecting the Arabic language into a small open-source model mainly trained in English. Our method demonstrates significant improvements in Arabic language performance, with an average 8% improvement across various benchmarks, while retaining the model's existing knowledge with a minimum amount of the original model's data. This offers a cost-effective alternative to training a comprehensive model in both English and Arabic. The results highlight the potential for efficient, targeted language model expansion without extensive retraining or resource-intensive processes.
format Preprint
id arxiv_https___arxiv_org_abs_2504_15120
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Kuwain 1.5B: An Arabic SLM via Language Injection
Hennara, Khalil
Chrouf, Sara
Hamed, Mohamed Motaism
Aldallal, Zeina
Hadid, Omar
AlModhayan, Safwan
Computation and Language
Artificial Intelligence
Enhancing existing models with new knowledge is a crucial aspect of AI development. This paper introduces a novel method for integrating a new language into a large language model (LLM). Our approach successfully incorporates a previously unseen target language into an existing LLM without compromising its prior knowledge. We trained a tiny model with 1.5 billion parameters named Kuwain by injecting the Arabic language into a small open-source model mainly trained in English. Our method demonstrates significant improvements in Arabic language performance, with an average 8% improvement across various benchmarks, while retaining the model's existing knowledge with a minimum amount of the original model's data. This offers a cost-effective alternative to training a comprehensive model in both English and Arabic. The results highlight the potential for efficient, targeted language model expansion without extensive retraining or resource-intensive processes.
title Kuwain 1.5B: An Arabic SLM via Language Injection
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2504.15120