Salvato in:
| Autori principali: | , , , , , , , , , , , , , , , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2407.12869 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866911968271532032 |
|---|---|
| author | Gosal, Gurpreet Xu, Yishi Ramakrishnan, Gokul Joshi, Rituraj Sheinin, Avraham Zhiming Chen Mishra, Biswajit Vassilieva, Natalia Hestness, Joel Sengupta, Neha Sahu, Sunil Kumar Jia, Bokang Pandit, Onkar Katipomu, Satheesh Kamboj, Samta Ghosh, Samujjwal Pal, Rahul Mullah, Parvez Doraiswamy, Soundar Chami, Mohamed El Karim Nakov, Preslav |
| author_facet | Gosal, Gurpreet Xu, Yishi Ramakrishnan, Gokul Joshi, Rituraj Sheinin, Avraham Zhiming Chen Mishra, Biswajit Vassilieva, Natalia Hestness, Joel Sengupta, Neha Sahu, Sunil Kumar Jia, Bokang Pandit, Onkar Katipomu, Satheesh Kamboj, Samta Ghosh, Samujjwal Pal, Rahul Mullah, Parvez Doraiswamy, Soundar Chami, Mohamed El Karim Nakov, Preslav |
| contents | We present an efficient method for adapting a monolingual Large Language Model (LLM) to another language, addressing challenges of catastrophic forgetting and tokenizer limitations. We focus this study on adapting Llama 2 to Arabic. Our two-stage approach begins with expanding the vocabulary and training only the embeddings matrix, followed by full model continual pre-training on a bilingual corpus. By continually pre-training on a mix of Arabic and English corpora, the model retains its proficiency in English while acquiring capabilities in Arabic. Our approach results in significant improvements in Arabic and slight enhancements in English, demonstrating cost-effective cross-lingual transfer. We perform ablations on embedding initialization techniques, data mix ratios, and learning rates and release a detailed training recipe. To demonstrate generalizability of this approach we also adapted Llama 3 8B to Arabic and Llama 2 13B to Hindi. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2407_12869 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Bilingual Adaptation of Monolingual Foundation Models Gosal, Gurpreet Xu, Yishi Ramakrishnan, Gokul Joshi, Rituraj Sheinin, Avraham Zhiming Chen Mishra, Biswajit Vassilieva, Natalia Hestness, Joel Sengupta, Neha Sahu, Sunil Kumar Jia, Bokang Pandit, Onkar Katipomu, Satheesh Kamboj, Samta Ghosh, Samujjwal Pal, Rahul Mullah, Parvez Doraiswamy, Soundar Chami, Mohamed El Karim Nakov, Preslav Computation and Language Artificial Intelligence We present an efficient method for adapting a monolingual Large Language Model (LLM) to another language, addressing challenges of catastrophic forgetting and tokenizer limitations. We focus this study on adapting Llama 2 to Arabic. Our two-stage approach begins with expanding the vocabulary and training only the embeddings matrix, followed by full model continual pre-training on a bilingual corpus. By continually pre-training on a mix of Arabic and English corpora, the model retains its proficiency in English while acquiring capabilities in Arabic. Our approach results in significant improvements in Arabic and slight enhancements in English, demonstrating cost-effective cross-lingual transfer. We perform ablations on embedding initialization techniques, data mix ratios, and learning rates and release a detailed training recipe. To demonstrate generalizability of this approach we also adapted Llama 3 8B to Arabic and Llama 2 13B to Hindi. |
| title | Bilingual Adaptation of Monolingual Foundation Models |
| topic | Computation and Language Artificial Intelligence |
| url | https://arxiv.org/abs/2407.12869 |