Theoretical Proof that Auto-regressive Language Models Collapse when Real-world Data is a Finite Set

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Wang, Lecheng, Shi, Xianjie, Li, Ge, Li, Jia, Zhang, Xuanming, Dong, Yihong, Jiao, Wenpin, Mei, Hong
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916743482441728
author Wang, Lecheng
Shi, Xianjie
Li, Ge
Li, Jia
Zhang, Xuanming
Dong, Yihong
Jiao, Wenpin
Mei, Hong
author_facet Wang, Lecheng
Shi, Xianjie
Li, Ge
Li, Jia
Zhang, Xuanming
Dong, Yihong
Jiao, Wenpin
Mei, Hong
contents Auto-regressive language models (LMs) have been widely used to generate data in data-scarce domains to train new LMs, compensating for the scarcity of real-world data. Previous work experimentally found that LMs collapse when trained on recursively generated data. This paper presents a theoretical proof: once a corpus (such as a subset of the World Wide Web) begins to incorporate generated data and no new real-world data is added to the corpus, then no matter how small the amount of data each LM generates and contributes to the corpus, LM collapse is inevitable after sufficient time. This finding suggests that attempts to mitigate collapse by limiting the quantity of synthetic data in the corpus are fundamentally insufficient. Instead, avoiding collapse hinges on ensuring the quality of synthetic data.
format Preprint
id arxiv_https___arxiv_org_abs_2412_14872
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Theoretical Proof that Auto-regressive Language Models Collapse when Real-world Data is a Finite Set
Wang, Lecheng
Shi, Xianjie
Li, Ge
Li, Jia
Zhang, Xuanming
Dong, Yihong
Jiao, Wenpin
Mei, Hong
Computation and Language
Auto-regressive language models (LMs) have been widely used to generate data in data-scarce domains to train new LMs, compensating for the scarcity of real-world data. Previous work experimentally found that LMs collapse when trained on recursively generated data. This paper presents a theoretical proof: once a corpus (such as a subset of the World Wide Web) begins to incorporate generated data and no new real-world data is added to the corpus, then no matter how small the amount of data each LM generates and contributes to the corpus, LM collapse is inevitable after sufficient time. This finding suggests that attempts to mitigate collapse by limiting the quantity of synthetic data in the corpus are fundamentally insufficient. Instead, avoiding collapse hinges on ensuring the quality of synthetic data.
title Theoretical Proof that Auto-regressive Language Models Collapse when Real-world Data is a Finite Set
topic Computation and Language
url https://arxiv.org/abs/2412.14872