Scaling Laws of Synthetic Data for Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Qin, Zeyu, Dong, Qingxiu, Zhang, Xingxing, Dong, Li, Huang, Xiaolong, Yang, Ziyi, Khademi, Mahmoud, Zhang, Dongdong, Awadalla, Hany Hassan, Fung, Yi R., Chen, Weizhu, Cheng, Minhao, Wei, Furu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909824922419200
author Qin, Zeyu
Dong, Qingxiu
Zhang, Xingxing
Dong, Li
Huang, Xiaolong
Yang, Ziyi
Khademi, Mahmoud
Zhang, Dongdong
Awadalla, Hany Hassan
Fung, Yi R.
Chen, Weizhu
Cheng, Minhao
Wei, Furu
author_facet Qin, Zeyu
Dong, Qingxiu
Zhang, Xingxing
Dong, Li
Huang, Xiaolong
Yang, Ziyi
Khademi, Mahmoud
Zhang, Dongdong
Awadalla, Hany Hassan
Fung, Yi R.
Chen, Weizhu
Cheng, Minhao
Wei, Furu
contents Large language models (LLMs) achieve strong performance across diverse tasks, largely driven by high-quality web data used in pre-training. However, recent studies indicate this data source is rapidly depleting. Synthetic data emerges as a promising alternative, but it remains unclear whether synthetic datasets exhibit predictable scalability comparable to raw pre-training data. In this work, we systematically investigate the scaling laws of synthetic data by introducing SynthLLM, a scalable framework that transforms pre-training corpora into diverse, high-quality synthetic datasets. Our approach achieves this by automatically extracting and recombining high-level concepts across multiple documents using a graph algorithm. Key findings from our extensive mathematical experiments on SynthLLM include: (1) SynthLLM generates synthetic data that reliably adheres to the rectified scaling law across various model sizes; (2) Performance improvements plateau near 300B tokens; and (3) Larger models approach optimal performance with fewer training tokens. For instance, an 8B model peaks at 1T tokens, while a 3B model requires 4T. Moreover, comparisons with existing synthetic data generation and augmentation methods demonstrate that SynthLLM achieves superior performance and scalability. Our findings highlight synthetic data as a scalable and reliable alternative to organic pre-training corpora, offering a viable path toward continued improvement in model performance.
format Preprint
id arxiv_https___arxiv_org_abs_2503_19551
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Scaling Laws of Synthetic Data for Language Models
Qin, Zeyu
Dong, Qingxiu
Zhang, Xingxing
Dong, Li
Huang, Xiaolong
Yang, Ziyi
Khademi, Mahmoud
Zhang, Dongdong
Awadalla, Hany Hassan
Fung, Yi R.
Chen, Weizhu
Cheng, Minhao
Wei, Furu
Computation and Language
Artificial Intelligence
Large language models (LLMs) achieve strong performance across diverse tasks, largely driven by high-quality web data used in pre-training. However, recent studies indicate this data source is rapidly depleting. Synthetic data emerges as a promising alternative, but it remains unclear whether synthetic datasets exhibit predictable scalability comparable to raw pre-training data. In this work, we systematically investigate the scaling laws of synthetic data by introducing SynthLLM, a scalable framework that transforms pre-training corpora into diverse, high-quality synthetic datasets. Our approach achieves this by automatically extracting and recombining high-level concepts across multiple documents using a graph algorithm. Key findings from our extensive mathematical experiments on SynthLLM include: (1) SynthLLM generates synthetic data that reliably adheres to the rectified scaling law across various model sizes; (2) Performance improvements plateau near 300B tokens; and (3) Larger models approach optimal performance with fewer training tokens. For instance, an 8B model peaks at 1T tokens, while a 3B model requires 4T. Moreover, comparisons with existing synthetic data generation and augmentation methods demonstrate that SynthLLM achieves superior performance and scalability. Our findings highlight synthetic data as a scalable and reliable alternative to organic pre-training corpora, offering a viable path toward continued improvement in model performance.
title Scaling Laws of Synthetic Data for Language Models
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2503.19551