The Automated LLM Speedrunning Benchmark: Reproducing NanoGPT Improvements

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhao, Bingchen, Magka, Despoina, Jiang, Minqi, Li, Xian, Raileanu, Roberta, Shavrina, Tatiana, Gagnon-Audet, Jean-Christophe, Niu, Kelvin, Sodhani, Shagun, Shvartsman, Michael, Lupu, Andrei, Lupidi, Alisia, Toledo, Edan, Hambardzumyan, Karen, Josifoski, Martin, Foster, Thomas, Cipolina-Kun, Lucia, Charnalia, Abhishek, Dunfield, Derek, Miller, Alexander H., Mac Aodha, Oisin, Foerster, Jakob, Bachrach, Yoram
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911030199713792
author Zhao, Bingchen
Magka, Despoina
Jiang, Minqi
Li, Xian
Raileanu, Roberta
Shavrina, Tatiana
Gagnon-Audet, Jean-Christophe
Niu, Kelvin
Sodhani, Shagun
Shvartsman, Michael
Lupu, Andrei
Lupidi, Alisia
Toledo, Edan
Hambardzumyan, Karen
Josifoski, Martin
Foster, Thomas
Cipolina-Kun, Lucia
Charnalia, Abhishek
Dunfield, Derek
Miller, Alexander H.
Mac Aodha, Oisin
Foerster, Jakob
Bachrach, Yoram
author_facet Zhao, Bingchen
Magka, Despoina
Jiang, Minqi
Li, Xian
Raileanu, Roberta
Shavrina, Tatiana
Gagnon-Audet, Jean-Christophe
Niu, Kelvin
Sodhani, Shagun
Shvartsman, Michael
Lupu, Andrei
Lupidi, Alisia
Toledo, Edan
Hambardzumyan, Karen
Josifoski, Martin
Foster, Thomas
Cipolina-Kun, Lucia
Charnalia, Abhishek
Dunfield, Derek
Miller, Alexander H.
Mac Aodha, Oisin
Foerster, Jakob
Bachrach, Yoram
contents Rapid advancements in large language models (LLMs) have the potential to assist in scientific progress. A critical capability toward this endeavor is the ability to reproduce existing work. To evaluate the ability of AI agents to reproduce results in an active research area, we introduce the Automated LLM Speedrunning Benchmark, leveraging the research community contributions on the NanoGPT speedrun, a competition to train a GPT-2 model in the shortest time. Each of the 19 speedrun tasks provides the agent with the previous records training script, optionally paired with one of three hint formats, ranging from pseudocode to paper-like descriptions of the new records improvements. Records execute quickly by design and speedrun improvements encompass diverse code-level changes, ranging from high-level algorithmic advancements to hardware-aware optimizations. These features make the benchmark both accessible and realistic for the frontier problem of improving LLM training. We find that recent reasoning LLMs combined with SoTA scaffolds struggle to reimplement already-known innovations in our benchmark, even when given detailed hints. Our benchmark thus provides a simple, non-saturated measure of an LLMs ability to automate scientific reproduction, a necessary (but not sufficient) skill for an autonomous research agent.
format Preprint
id arxiv_https___arxiv_org_abs_2506_22419
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle The Automated LLM Speedrunning Benchmark: Reproducing NanoGPT Improvements
Zhao, Bingchen
Magka, Despoina
Jiang, Minqi
Li, Xian
Raileanu, Roberta
Shavrina, Tatiana
Gagnon-Audet, Jean-Christophe
Niu, Kelvin
Sodhani, Shagun
Shvartsman, Michael
Lupu, Andrei
Lupidi, Alisia
Toledo, Edan
Hambardzumyan, Karen
Josifoski, Martin
Foster, Thomas
Cipolina-Kun, Lucia
Charnalia, Abhishek
Dunfield, Derek
Miller, Alexander H.
Mac Aodha, Oisin
Foerster, Jakob
Bachrach, Yoram
Artificial Intelligence
Computation and Language
Machine Learning
Rapid advancements in large language models (LLMs) have the potential to assist in scientific progress. A critical capability toward this endeavor is the ability to reproduce existing work. To evaluate the ability of AI agents to reproduce results in an active research area, we introduce the Automated LLM Speedrunning Benchmark, leveraging the research community contributions on the NanoGPT speedrun, a competition to train a GPT-2 model in the shortest time. Each of the 19 speedrun tasks provides the agent with the previous records training script, optionally paired with one of three hint formats, ranging from pseudocode to paper-like descriptions of the new records improvements. Records execute quickly by design and speedrun improvements encompass diverse code-level changes, ranging from high-level algorithmic advancements to hardware-aware optimizations. These features make the benchmark both accessible and realistic for the frontier problem of improving LLM training. We find that recent reasoning LLMs combined with SoTA scaffolds struggle to reimplement already-known innovations in our benchmark, even when given detailed hints. Our benchmark thus provides a simple, non-saturated measure of an LLMs ability to automate scientific reproduction, a necessary (but not sufficient) skill for an autonomous research agent.
title The Automated LLM Speedrunning Benchmark: Reproducing NanoGPT Improvements
topic Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2506.22419