Salvato in:
Dettagli Bibliografici
Autori principali: Gosal, Gurpreet, Xu, Yishi, Ramakrishnan, Gokul, Joshi, Rituraj, Sheinin, Avraham, Zhiming, Chen, Mishra, Biswajit, Vassilieva, Natalia, Hestness, Joel, Sengupta, Neha, Sahu, Sunil Kumar, Jia, Bokang, Pandit, Onkar, Katipomu, Satheesh, Kamboj, Samta, Ghosh, Samujjwal, Pal, Rahul, Mullah, Parvez, Doraiswamy, Soundar, Chami, Mohamed El Karim, Nakov, Preslav
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:https://arxiv.org/abs/2407.12869
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911968271532032
author Gosal, Gurpreet
Xu, Yishi
Ramakrishnan, Gokul
Joshi, Rituraj
Sheinin, Avraham
Zhiming
Chen
Mishra, Biswajit
Vassilieva, Natalia
Hestness, Joel
Sengupta, Neha
Sahu, Sunil Kumar
Jia, Bokang
Pandit, Onkar
Katipomu, Satheesh
Kamboj, Samta
Ghosh, Samujjwal
Pal, Rahul
Mullah, Parvez
Doraiswamy, Soundar
Chami, Mohamed El Karim
Nakov, Preslav
author_facet Gosal, Gurpreet
Xu, Yishi
Ramakrishnan, Gokul
Joshi, Rituraj
Sheinin, Avraham
Zhiming
Chen
Mishra, Biswajit
Vassilieva, Natalia
Hestness, Joel
Sengupta, Neha
Sahu, Sunil Kumar
Jia, Bokang
Pandit, Onkar
Katipomu, Satheesh
Kamboj, Samta
Ghosh, Samujjwal
Pal, Rahul
Mullah, Parvez
Doraiswamy, Soundar
Chami, Mohamed El Karim
Nakov, Preslav
contents We present an efficient method for adapting a monolingual Large Language Model (LLM) to another language, addressing challenges of catastrophic forgetting and tokenizer limitations. We focus this study on adapting Llama 2 to Arabic. Our two-stage approach begins with expanding the vocabulary and training only the embeddings matrix, followed by full model continual pre-training on a bilingual corpus. By continually pre-training on a mix of Arabic and English corpora, the model retains its proficiency in English while acquiring capabilities in Arabic. Our approach results in significant improvements in Arabic and slight enhancements in English, demonstrating cost-effective cross-lingual transfer. We perform ablations on embedding initialization techniques, data mix ratios, and learning rates and release a detailed training recipe. To demonstrate generalizability of this approach we also adapted Llama 3 8B to Arabic and Llama 2 13B to Hindi.
format Preprint
id arxiv_https___arxiv_org_abs_2407_12869
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Bilingual Adaptation of Monolingual Foundation Models
Gosal, Gurpreet
Xu, Yishi
Ramakrishnan, Gokul
Joshi, Rituraj
Sheinin, Avraham
Zhiming
Chen
Mishra, Biswajit
Vassilieva, Natalia
Hestness, Joel
Sengupta, Neha
Sahu, Sunil Kumar
Jia, Bokang
Pandit, Onkar
Katipomu, Satheesh
Kamboj, Samta
Ghosh, Samujjwal
Pal, Rahul
Mullah, Parvez
Doraiswamy, Soundar
Chami, Mohamed El Karim
Nakov, Preslav
Computation and Language
Artificial Intelligence
We present an efficient method for adapting a monolingual Large Language Model (LLM) to another language, addressing challenges of catastrophic forgetting and tokenizer limitations. We focus this study on adapting Llama 2 to Arabic. Our two-stage approach begins with expanding the vocabulary and training only the embeddings matrix, followed by full model continual pre-training on a bilingual corpus. By continually pre-training on a mix of Arabic and English corpora, the model retains its proficiency in English while acquiring capabilities in Arabic. Our approach results in significant improvements in Arabic and slight enhancements in English, demonstrating cost-effective cross-lingual transfer. We perform ablations on embedding initialization techniques, data mix ratios, and learning rates and release a detailed training recipe. To demonstrate generalizability of this approach we also adapted Llama 3 8B to Arabic and Llama 2 13B to Hindi.
title Bilingual Adaptation of Monolingual Foundation Models
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2407.12869