Rephrasing natural text data with different languages and quality levels for Large Language Model pre-training

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Pieler, Michael, Bellagente, Marco, Teufel, Hannah, Phung, Duy, Cooper, Nathan, Tow, Jonathan, Rocha, Paulo, Adithyan, Reshinth, Alyafeai, Zaid, Pinnaparaju, Nikhil, Zhuravinskyi, Maksym, Riquelme, Carlos
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909367898472448
author Pieler, Michael
Bellagente, Marco
Teufel, Hannah
Phung, Duy
Cooper, Nathan
Tow, Jonathan
Rocha, Paulo
Adithyan, Reshinth
Alyafeai, Zaid
Pinnaparaju, Nikhil
Zhuravinskyi, Maksym
Riquelme, Carlos
author_facet Pieler, Michael
Bellagente, Marco
Teufel, Hannah
Phung, Duy
Cooper, Nathan
Tow, Jonathan
Rocha, Paulo
Adithyan, Reshinth
Alyafeai, Zaid
Pinnaparaju, Nikhil
Zhuravinskyi, Maksym
Riquelme, Carlos
contents Recently published work on rephrasing natural text data for pre-training LLMs has shown promising results when combining the original dataset with the synthetically rephrased data. We build upon previous work by replicating existing results on C4 and extending them with our optimized rephrasing pipeline to the English, German, Italian, and Spanish Oscar subsets of CulturaX. Our pipeline leads to increased performance on standard evaluation benchmarks in both the mono- and multilingual setup. In addition, we provide a detailed study of our pipeline, investigating the choice of the base dataset and LLM for the rephrasing, as well as the relationship between the model size and the performance after pre-training. By exploring data with different perceived quality levels, we show that gains decrease with higher quality. Furthermore, we find the difference in performance between model families to be bigger than between different model sizes. This highlights the necessity for detailed tests before choosing an LLM to rephrase large amounts of data. Moreover, we investigate the effect of pre-training with synthetic data on supervised fine-tuning. Here, we find increasing but inconclusive results that highly depend on the used benchmark. These results (again) highlight the need for better benchmarking setups. In summary, we show that rephrasing multilingual and low-quality data is a very promising direction to extend LLM pre-training data.
format Preprint
id arxiv_https___arxiv_org_abs_2410_20796
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Rephrasing natural text data with different languages and quality levels for Large Language Model pre-training
Pieler, Michael
Bellagente, Marco
Teufel, Hannah
Phung, Duy
Cooper, Nathan
Tow, Jonathan
Rocha, Paulo
Adithyan, Reshinth
Alyafeai, Zaid
Pinnaparaju, Nikhil
Zhuravinskyi, Maksym
Riquelme, Carlos
Computation and Language
Recently published work on rephrasing natural text data for pre-training LLMs has shown promising results when combining the original dataset with the synthetically rephrased data. We build upon previous work by replicating existing results on C4 and extending them with our optimized rephrasing pipeline to the English, German, Italian, and Spanish Oscar subsets of CulturaX. Our pipeline leads to increased performance on standard evaluation benchmarks in both the mono- and multilingual setup. In addition, we provide a detailed study of our pipeline, investigating the choice of the base dataset and LLM for the rephrasing, as well as the relationship between the model size and the performance after pre-training. By exploring data with different perceived quality levels, we show that gains decrease with higher quality. Furthermore, we find the difference in performance between model families to be bigger than between different model sizes. This highlights the necessity for detailed tests before choosing an LLM to rephrase large amounts of data. Moreover, we investigate the effect of pre-training with synthetic data on supervised fine-tuning. Here, we find increasing but inconclusive results that highly depend on the used benchmark. These results (again) highlight the need for better benchmarking setups. In summary, we show that rephrasing multilingual and low-quality data is a very promising direction to extend LLM pre-training data.
title Rephrasing natural text data with different languages and quality levels for Large Language Model pre-training
topic Computation and Language
url https://arxiv.org/abs/2410.20796