Multilingual Pretraining Using a Large Corpus Machine-Translated from a Single Source Language

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Wang, Jiayi, Lu, Yao, Weber, Maurice, Ryabinin, Max, Chen, Yihong, Tang, Raphael, Stenetorp, Pontus
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912106734944256
author Wang, Jiayi
Lu, Yao
Weber, Maurice
Ryabinin, Max
Chen, Yihong
Tang, Raphael
Stenetorp, Pontus
author_facet Wang, Jiayi
Lu, Yao
Weber, Maurice
Ryabinin, Max
Chen, Yihong
Tang, Raphael
Stenetorp, Pontus
contents English, as a very high-resource language, enables the pretraining of high-quality large language models (LLMs). The same cannot be said for most other languages, as leading LLMs still underperform for non-English languages, likely due to a gap in the quality and diversity of the available multilingual pretraining corpora. In this work, we find that machine-translated text from a single high-quality source language can contribute significantly to the pretraining of multilingual LLMs. We translate FineWeb-Edu, a high-quality English web dataset, into French, German, and Spanish, resulting in a final 300B-token dataset, which we call TransWeb-Edu, and pretrain a 1.3B-parameter model, CuatroLLM, from scratch on this dataset. Across five non-English reasoning tasks, we show that CuatroLLM matches or outperforms state-of-the-art multilingual models trained using closed data, such as Llama3.2 and Gemma2, despite using an order of magnitude less data, such as about 6% of the tokens used for Llama3.2's training. We further demonstrate that with additional domain-specific pretraining, amounting to less than 1% of TransWeb-Edu, CuatroLLM surpasses the state of the art in multilingual reasoning. To promote reproducibility, we release our corpus, models, and training pipeline under open licenses at hf.co/britllm/CuatroLLM.
format Preprint
id arxiv_https___arxiv_org_abs_2410_23956
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Multilingual Pretraining Using a Large Corpus Machine-Translated from a Single Source Language
Wang, Jiayi
Lu, Yao
Weber, Maurice
Ryabinin, Max
Chen, Yihong
Tang, Raphael
Stenetorp, Pontus
Computation and Language
English, as a very high-resource language, enables the pretraining of high-quality large language models (LLMs). The same cannot be said for most other languages, as leading LLMs still underperform for non-English languages, likely due to a gap in the quality and diversity of the available multilingual pretraining corpora. In this work, we find that machine-translated text from a single high-quality source language can contribute significantly to the pretraining of multilingual LLMs. We translate FineWeb-Edu, a high-quality English web dataset, into French, German, and Spanish, resulting in a final 300B-token dataset, which we call TransWeb-Edu, and pretrain a 1.3B-parameter model, CuatroLLM, from scratch on this dataset. Across five non-English reasoning tasks, we show that CuatroLLM matches or outperforms state-of-the-art multilingual models trained using closed data, such as Llama3.2 and Gemma2, despite using an order of magnitude less data, such as about 6% of the tokens used for Llama3.2's training. We further demonstrate that with additional domain-specific pretraining, amounting to less than 1% of TransWeb-Edu, CuatroLLM surpasses the state of the art in multilingual reasoning. To promote reproducibility, we release our corpus, models, and training pipeline under open licenses at hf.co/britllm/CuatroLLM.
title Multilingual Pretraining Using a Large Corpus Machine-Translated from a Single Source Language
topic Computation and Language
url https://arxiv.org/abs/2410.23956