Adapting Multilingual LLMs to Low-Resource Languages using Continued Pre-training and Synthetic Corpus

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Joshi, Raviraj, Singla, Kanishk, Kamath, Anusha, Kalani, Raunak, Paul, Rakesh, Vaidya, Utkarsh, Chauhan, Sanjay Singh, Wartikar, Niranjan, Long, Eileen
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912335996649472
author Joshi, Raviraj
Singla, Kanishk
Kamath, Anusha
Kalani, Raunak
Paul, Rakesh
Vaidya, Utkarsh
Chauhan, Sanjay Singh
Wartikar, Niranjan
Long, Eileen
author_facet Joshi, Raviraj
Singla, Kanishk
Kamath, Anusha
Kalani, Raunak
Paul, Rakesh
Vaidya, Utkarsh
Chauhan, Sanjay Singh
Wartikar, Niranjan
Long, Eileen
contents Multilingual LLMs support a variety of languages; however, their performance is suboptimal for low-resource languages. In this work, we emphasize the importance of continued pre-training of multilingual LLMs and the use of translation-based synthetic pre-training corpora for improving LLMs in low-resource languages. We conduct our study in the context of the low-resource Indic language Hindi. We introduce Nemotron-Mini-Hindi 4B, a bilingual SLM supporting both Hindi and English, based on Nemotron-Mini 4B. The model is trained using a mix of real and synthetic Hindi + English tokens, with continuous pre-training performed on 400B tokens. We demonstrate that both the base and instruct models achieve state-of-the-art results on Hindi benchmarks while remaining competitive on English tasks. Additionally, we observe that the continued pre-training approach enhances the model's overall factual accuracy. We perform an ablation study to highlight the impact of Hindi pre-training, showing significant improvements in Hindi chat capabilities and factual accuracy, which cannot be achieved through Hindi alignment alone.
format Preprint
id arxiv_https___arxiv_org_abs_2410_14815
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Adapting Multilingual LLMs to Low-Resource Languages using Continued Pre-training and Synthetic Corpus
Joshi, Raviraj
Singla, Kanishk
Kamath, Anusha
Kalani, Raunak
Paul, Rakesh
Vaidya, Utkarsh
Chauhan, Sanjay Singh
Wartikar, Niranjan
Long, Eileen
Computation and Language
Machine Learning
Multilingual LLMs support a variety of languages; however, their performance is suboptimal for low-resource languages. In this work, we emphasize the importance of continued pre-training of multilingual LLMs and the use of translation-based synthetic pre-training corpora for improving LLMs in low-resource languages. We conduct our study in the context of the low-resource Indic language Hindi. We introduce Nemotron-Mini-Hindi 4B, a bilingual SLM supporting both Hindi and English, based on Nemotron-Mini 4B. The model is trained using a mix of real and synthetic Hindi + English tokens, with continuous pre-training performed on 400B tokens. We demonstrate that both the base and instruct models achieve state-of-the-art results on Hindi benchmarks while remaining competitive on English tasks. Additionally, we observe that the continued pre-training approach enhances the model's overall factual accuracy. We perform an ablation study to highlight the impact of Hindi pre-training, showing significant improvements in Hindi chat capabilities and factual accuracy, which cannot be achieved through Hindi alignment alone.
title Adapting Multilingual LLMs to Low-Resource Languages using Continued Pre-training and Synthetic Corpus
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2410.14815