Training Bilingual LMs with Data Constraints in the Targeted Language

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Seto, Skyler, ter Hoeve, Maartje, Bai, Richard He, Schluter, Natalie, Grangier, David
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866913680743989248
author Seto, Skyler
ter Hoeve, Maartje
Bai, Richard He
Schluter, Natalie
Grangier, David
author_facet Seto, Skyler
ter Hoeve, Maartje
Bai, Richard He
Schluter, Natalie
Grangier, David
contents Large language models are trained on massive scrapes of the web, as required by current scaling laws. Most progress is made for English, given its abundance of high-quality pretraining data. For most other languages, however, such high quality pretraining data is unavailable. In this work, we study how to boost pretrained model performance in a target language with insufficient pretraining data for training a high performing language model, by enlisting data from an auxiliary language for which high quality data is available. We study this by quantifying the performance gap between training with data in a data-rich auxiliary language compared with training in the target language, exploring the benefits of translation systems, studying the limitations of model scaling when data is limited in the target languages, and proposing new methods for upsampling data from the auxiliary language. Our results show that stronger auxiliary datasets result in performance gains without modification to the model or training objective for close languages, and, in particular, that performance gains due to the development of more information-rich English pretraining datasets can extend to targeted language settings with limited data.
format Preprint
id arxiv_https___arxiv_org_abs_2411_12986
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Training Bilingual LMs with Data Constraints in the Targeted Language
Seto, Skyler
ter Hoeve, Maartje
Bai, Richard He
Schluter, Natalie
Grangier, David
Computation and Language
Machine Learning
Large language models are trained on massive scrapes of the web, as required by current scaling laws. Most progress is made for English, given its abundance of high-quality pretraining data. For most other languages, however, such high quality pretraining data is unavailable. In this work, we study how to boost pretrained model performance in a target language with insufficient pretraining data for training a high performing language model, by enlisting data from an auxiliary language for which high quality data is available. We study this by quantifying the performance gap between training with data in a data-rich auxiliary language compared with training in the target language, exploring the benefits of translation systems, studying the limitations of model scaling when data is limited in the target languages, and proposing new methods for upsampling data from the auxiliary language. Our results show that stronger auxiliary datasets result in performance gains without modification to the model or training objective for close languages, and, in particular, that performance gains due to the development of more information-rich English pretraining datasets can extend to targeted language settings with limited data.
title Training Bilingual LMs with Data Constraints in the Targeted Language
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2411.12986