Algorithmic progress in language models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Ho, Anson, Besiroglu, Tamay, Erdil, Ege, Owen, David, Rahman, Robi, Guo, Zifan Carl, Atkinson, David, Thompson, Neil, Sevilla, Jaime
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916155015299072
author Ho, Anson
Besiroglu, Tamay
Erdil, Ege
Owen, David
Rahman, Robi
Guo, Zifan Carl
Atkinson, David
Thompson, Neil
Sevilla, Jaime
author_facet Ho, Anson
Besiroglu, Tamay
Erdil, Ege
Owen, David
Rahman, Robi
Guo, Zifan Carl
Atkinson, David
Thompson, Neil
Sevilla, Jaime
contents We investigate the rate at which algorithms for pre-training language models have improved since the advent of deep learning. Using a dataset of over 200 language model evaluations on Wikitext and Penn Treebank spanning 2012-2023, we find that the compute required to reach a set performance threshold has halved approximately every 8 months, with a 95% confidence interval of around 5 to 14 months, substantially faster than hardware gains per Moore's Law. We estimate augmented scaling laws, which enable us to quantify algorithmic progress and determine the relative contributions of scaling models versus innovations in training algorithms. Despite the rapid pace of algorithmic progress and the development of new architectures such as the transformer, our analysis reveals that the increase in compute made an even larger contribution to overall performance improvements over this time period. Though limited by noisy benchmark data, our analysis quantifies the rapid progress in language modeling, shedding light on the relative contributions from compute and algorithms.
format Preprint
id arxiv_https___arxiv_org_abs_2403_05812
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Algorithmic progress in language models
Ho, Anson
Besiroglu, Tamay
Erdil, Ege
Owen, David
Rahman, Robi
Guo, Zifan Carl
Atkinson, David
Thompson, Neil
Sevilla, Jaime
Computation and Language
Artificial Intelligence
We investigate the rate at which algorithms for pre-training language models have improved since the advent of deep learning. Using a dataset of over 200 language model evaluations on Wikitext and Penn Treebank spanning 2012-2023, we find that the compute required to reach a set performance threshold has halved approximately every 8 months, with a 95% confidence interval of around 5 to 14 months, substantially faster than hardware gains per Moore's Law. We estimate augmented scaling laws, which enable us to quantify algorithmic progress and determine the relative contributions of scaling models versus innovations in training algorithms. Despite the rapid pace of algorithmic progress and the development of new architectures such as the transformer, our analysis reveals that the increase in compute made an even larger contribution to overall performance improvements over this time period. Though limited by noisy benchmark data, our analysis quantifies the rapid progress in language modeling, shedding light on the relative contributions from compute and algorithms.
title Algorithmic progress in language models
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2403.05812