Kuwain 1.5B: An Arabic SLM via Language Injection
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | , , , , , |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866915453752836096 |
|---|---|
| author | Hennara, Khalil Chrouf, Sara Hamed, Mohamed Motaism Aldallal, Zeina Hadid, Omar AlModhayan, Safwan |
| author_facet | Hennara, Khalil Chrouf, Sara Hamed, Mohamed Motaism Aldallal, Zeina Hadid, Omar AlModhayan, Safwan |
| contents | Enhancing existing models with new knowledge is a crucial aspect of AI development. This paper introduces a novel method for integrating a new language into a large language model (LLM). Our approach successfully incorporates a previously unseen target language into an existing LLM without compromising its prior knowledge. We trained a tiny model with 1.5 billion parameters named Kuwain by injecting the Arabic language into a small open-source model mainly trained in English. Our method demonstrates significant improvements in Arabic language performance, with an average 8% improvement across various benchmarks, while retaining the model's existing knowledge with a minimum amount of the original model's data. This offers a cost-effective alternative to training a comprehensive model in both English and Arabic. The results highlight the potential for efficient, targeted language model expansion without extensive retraining or resource-intensive processes. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2504_15120 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Kuwain 1.5B: An Arabic SLM via Language Injection Hennara, Khalil Chrouf, Sara Hamed, Mohamed Motaism Aldallal, Zeina Hadid, Omar AlModhayan, Safwan Computation and Language Artificial Intelligence Enhancing existing models with new knowledge is a crucial aspect of AI development. This paper introduces a novel method for integrating a new language into a large language model (LLM). Our approach successfully incorporates a previously unseen target language into an existing LLM without compromising its prior knowledge. We trained a tiny model with 1.5 billion parameters named Kuwain by injecting the Arabic language into a small open-source model mainly trained in English. Our method demonstrates significant improvements in Arabic language performance, with an average 8% improvement across various benchmarks, while retaining the model's existing knowledge with a minimum amount of the original model's data. This offers a cost-effective alternative to training a comprehensive model in both English and Arabic. The results highlight the potential for efficient, targeted language model expansion without extensive retraining or resource-intensive processes. |
| title | Kuwain 1.5B: An Arabic SLM via Language Injection |
| topic | Computation and Language Artificial Intelligence |
| url | https://arxiv.org/abs/2504.15120 |