How to Synthesize Text Data without Model Collapse?

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Zhu, Xuekai, Cheng, Daixuan, Li, Hengli, Zhang, Kaiyan, Hua, Ermo, Lv, Xingtai, Ding, Ning, Lin, Zhouhan, Zheng, Zilong, Zhou, Bowen
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866913862723305472
author Zhu, Xuekai
Cheng, Daixuan
Li, Hengli
Zhang, Kaiyan
Hua, Ermo
Lv, Xingtai
Ding, Ning
Lin, Zhouhan
Zheng, Zilong
Zhou, Bowen
author_facet Zhu, Xuekai
Cheng, Daixuan
Li, Hengli
Zhang, Kaiyan
Hua, Ermo
Lv, Xingtai
Ding, Ning
Lin, Zhouhan
Zheng, Zilong
Zhou, Bowen
contents Model collapse in synthetic data indicates that iterative training on self-generated data leads to a gradual decline in performance. With the proliferation of AI models, synthetic data will fundamentally reshape the web data ecosystem. Future GPT-$\{n\}$ models will inevitably be trained on a blend of synthetic and human-produced data. In this paper, we focus on two questions: what is the impact of synthetic data on language model training, and how to synthesize data without model collapse? We first pre-train language models across different proportions of synthetic data, revealing a negative correlation between the proportion of synthetic data and model performance. We further conduct statistical analysis on synthetic data to uncover distributional shift phenomenon and over-concentration of n-gram features. Inspired by the above findings, we propose token editing on human-produced data to obtain semi-synthetic data. As a proof of concept, we theoretically demonstrate that token-level editing can prevent model collapse, as the test error is constrained by a finite upper bound. We conduct extensive experiments on pre-training from scratch, continual pre-training, and supervised fine-tuning. The results validate our theoretical proof that token-level editing improves model performance.
format Preprint
id arxiv_https___arxiv_org_abs_2412_14689
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle How to Synthesize Text Data without Model Collapse?
Zhu, Xuekai
Cheng, Daixuan
Li, Hengli
Zhang, Kaiyan
Hua, Ermo
Lv, Xingtai
Ding, Ning
Lin, Zhouhan
Zheng, Zilong
Zhou, Bowen
Computation and Language
Artificial Intelligence
Machine Learning
Model collapse in synthetic data indicates that iterative training on self-generated data leads to a gradual decline in performance. With the proliferation of AI models, synthetic data will fundamentally reshape the web data ecosystem. Future GPT-$\{n\}$ models will inevitably be trained on a blend of synthetic and human-produced data. In this paper, we focus on two questions: what is the impact of synthetic data on language model training, and how to synthesize data without model collapse? We first pre-train language models across different proportions of synthetic data, revealing a negative correlation between the proportion of synthetic data and model performance. We further conduct statistical analysis on synthetic data to uncover distributional shift phenomenon and over-concentration of n-gram features. Inspired by the above findings, we propose token editing on human-produced data to obtain semi-synthetic data. As a proof of concept, we theoretically demonstrate that token-level editing can prevent model collapse, as the test error is constrained by a finite upper bound. We conduct extensive experiments on pre-training from scratch, continual pre-training, and supervised fine-tuning. The results validate our theoretical proof that token-level editing improves model performance.
title How to Synthesize Text Data without Model Collapse?
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2412.14689