Will we run out of data? Limits of LLM scaling based on human-generated data

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Villalobos, Pablo, Ho, Anson, Sevilla, Jaime, Besiroglu, Tamay, Heim, Lennart, Hobbhahn, Marius
Format: Preprint
Published: 2022
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916274003509248
author Villalobos, Pablo
Ho, Anson
Sevilla, Jaime
Besiroglu, Tamay
Heim, Lennart
Hobbhahn, Marius
author_facet Villalobos, Pablo
Ho, Anson
Sevilla, Jaime
Besiroglu, Tamay
Heim, Lennart
Hobbhahn, Marius
contents We investigate the potential constraints on LLM scaling posed by the availability of public human-generated text data. We forecast the growing demand for training data based on current trends and estimate the total stock of public human text data. Our findings indicate that if current LLM development trends continue, models will be trained on datasets roughly equal in size to the available stock of public human text data between 2026 and 2032, or slightly earlier if models are overtrained. We explore how progress in language modeling can continue when human-generated text datasets cannot be scaled any further. We argue that synthetic data generation, transfer learning from data-rich domains, and data efficiency improvements might support further progress.
format Preprint
id arxiv_https___arxiv_org_abs_2211_04325
institution arXiv
publishDate 2022
record_format arxiv
spellingShingle Will we run out of data? Limits of LLM scaling based on human-generated data
Villalobos, Pablo
Ho, Anson
Sevilla, Jaime
Besiroglu, Tamay
Heim, Lennart
Hobbhahn, Marius
Machine Learning
Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
Computers and Society
We investigate the potential constraints on LLM scaling posed by the availability of public human-generated text data. We forecast the growing demand for training data based on current trends and estimate the total stock of public human text data. Our findings indicate that if current LLM development trends continue, models will be trained on datasets roughly equal in size to the available stock of public human text data between 2026 and 2032, or slightly earlier if models are overtrained. We explore how progress in language modeling can continue when human-generated text datasets cannot be scaled any further. We argue that synthetic data generation, transfer learning from data-rich domains, and data efficiency improvements might support further progress.
title Will we run out of data? Limits of LLM scaling based on human-generated data
topic Machine Learning
Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
Computers and Society
url https://arxiv.org/abs/2211.04325