The Automated LLM Speedrunning Benchmark: Reproducing NanoGPT Improvements
Fuente:
arXiv
Saved in:
| Main Authors: | Zhao, Bingchen, Magka, Despoina, Jiang, Minqi, Li, Xian, Raileanu, Roberta, Shavrina, Tatiana, Gagnon-Audet, Jean-Christophe, Niu, Kelvin, Sodhani, Shagun, Shvartsman, Michael, Lupu, Andrei, Lupidi, Alisia, Toledo, Edan, Hambardzumyan, Karen, Josifoski, Martin, Foster, Thomas, Cipolina-Kun, Lucia, Charnalia, Abhishek, Dunfield, Derek, Miller, Alexander H., Mac Aodha, Oisin, Foerster, Jakob, Bachrach, Yoram |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
APRES: An Agentic Paper Revision and Evaluation System
by: Zhao, Bingchen, et al.
Published: (2026)
by: Zhao, Bingchen, et al.
Published: (2026)
AI Research Agents for Machine Learning: Search, Exploration, and Generalization in MLE-bench
by: Toledo, Edan, et al.
Published: (2025)
by: Toledo, Edan, et al.
Published: (2025)
Bootstrapping Task Spaces for Self-Improvement
by: Jiang, Minqi, et al.
Published: (2025)
by: Jiang, Minqi, et al.
Published: (2025)
What Does It Take to Be a Good AI Research Agent? Studying the Role of Ideation Diversity
by: Audran-Reiss, Alexis, et al.
Published: (2025)
by: Audran-Reiss, Alexis, et al.
Published: (2025)
AIRA_2: Overcoming Bottlenecks in AI Research Agents
by: Hambardzumyan, Karen, et al.
Published: (2026)
by: Hambardzumyan, Karen, et al.
Published: (2026)
AIRS-Bench: a Suite of Tasks for Frontier AI Research Science Agents
by: Lupidi, Alisia, et al.
Published: (2026)
by: Lupidi, Alisia, et al.
Published: (2026)
Souper-Model: How Simple Arithmetic Unlocks State-of-the-Art LLM Performance
by: Maiti, Shalini, et al.
Published: (2025)
by: Maiti, Shalini, et al.
Published: (2025)
Interpretable Text-Guided Image Clustering via Iterative Search
by: Zhao, Bingchen, et al.
Published: (2025)
by: Zhao, Bingchen, et al.
Published: (2025)
Scaling and Distilling Transformer Models for sEMG
by: Mehlman, Nicholas, et al.
Published: (2025)
by: Mehlman, Nicholas, et al.
Published: (2025)
Labeled Data Selection for Category Discovery
by: Zhao, Bingchen, et al.
Published: (2024)
by: Zhao, Bingchen, et al.
Published: (2024)
MLGym: A New Framework and Benchmark for Advancing AI Research Agents
by: Nathani, Deepak, et al.
Published: (2025)
by: Nathani, Deepak, et al.
Published: (2025)
Agentic Discovery of Neural Architectures: AIRA-Compose and AIRA-Design
by: Pepe, Alberto, et al.
Published: (2026)
by: Pepe, Alberto, et al.
Published: (2026)
Source2Synth: Synthetic Data Generation and Curation Grounded in Real Data Sources
by: Lupidi, Alisia, et al.
Published: (2024)
by: Lupidi, Alisia, et al.
Published: (2024)
SAOR: Single-View Articulated Object Reconstruction
by: Aygün, Mehmet, et al.
Published: (2023)
by: Aygün, Mehmet, et al.
Published: (2023)
Effect of Document Packing on the Latent Multi-Hop Reasoning Capabilities of Large Language Models
by: Prato, Gabriele, et al.
Published: (2025)
by: Prato, Gabriele, et al.
Published: (2025)
Vision Learners Meet Web Image-Text Pairs
by: Zhao, Bingchen, et al.
Published: (2023)
by: Zhao, Bingchen, et al.
Published: (2023)
Asking the Right Questions: Improving Reasoning with Generated Stepping Stones
by: Hu, Hengyuan, et al.
Published: (2026)
by: Hu, Hengyuan, et al.
Published: (2026)
Speedrunning and path integrals
by: Lami, Gabriele
Published: (2024)
by: Lami, Gabriele
Published: (2024)
Hyperagents
by: Zhang, Jenny, et al.
Published: (2026)
by: Zhang, Jenny, et al.
Published: (2026)
The Generalization Gap in Offline Reinforcement Learning
by: Mediratta, Ishita, et al.
Published: (2023)
by: Mediratta, Ishita, et al.
Published: (2023)
Speedrunning ImageNet Diffusion
by: Bhanded, Swayam
Published: (2025)
by: Bhanded, Swayam
Published: (2025)
Self-Supervised Multimodal Learning: A Survey
by: Zong, Yongshuo, et al.
Published: (2023)
by: Zong, Yongshuo, et al.
Published: (2023)
Improving Semantic Correspondence with Viewpoint-Guided Spherical Maps
by: Mariotti, Octave, et al.
Published: (2023)
by: Mariotti, Octave, et al.
Published: (2023)
Representational Similarity via Interpretable Visual Concepts
by: Kondapaneni, Neehar, et al.
Published: (2025)
by: Kondapaneni, Neehar, et al.
Published: (2025)
Representational Difference Explanations
by: Kondapaneni, Neehar, et al.
Published: (2025)
by: Kondapaneni, Neehar, et al.
Published: (2025)
When should we prefer Decision Transformers for Offline Reinforcement Learning?
by: Bhargava, Prajjwal, et al.
Published: (2023)
by: Bhargava, Prajjwal, et al.
Published: (2023)
Do Large Language Models Know How Much They Know?
by: Prato, Gabriele, et al.
Published: (2025)
by: Prato, Gabriele, et al.
Published: (2025)
EpiK-Eval: Evaluation for Language Models as Epistemic Models
by: Prato, Gabriele, et al.
Published: (2023)
by: Prato, Gabriele, et al.
Published: (2023)
Harnessing small projectors and multiple views for efficient vision pretraining
by: Agrawal, Kumar Krishna, et al.
Published: (2023)
by: Agrawal, Kumar Krishna, et al.
Published: (2023)
Neural Mean-Field Games: Extending Mean-Field Game Theory with Neural Stochastic Differential Equations
by: Thöni, Anna C. M., et al.
Published: (2025)
by: Thöni, Anna C. M., et al.
Published: (2025)
MaestroMotif: Skill Design from Artificial Intelligence Feedback
by: Klissarov, Martin, et al.
Published: (2024)
by: Klissarov, Martin, et al.
Published: (2024)
Scaling Small Agents Through Strategy Auctions
by: Alazraki, Lisa, et al.
Published: (2026)
by: Alazraki, Lisa, et al.
Published: (2026)
Less is More: Discovering Concise Network Explanations
by: Kondapaneni, Neehar, et al.
Published: (2024)
by: Kondapaneni, Neehar, et al.
Published: (2024)
Generating Binary Species Range Maps
by: Dorm, Filip, et al.
Published: (2024)
by: Dorm, Filip, et al.
Published: (2024)
MotionPhysics: Learnable Motion Distillation for Text-Guided Simulation
by: Wang, Miaowei, et al.
Published: (2026)
by: Wang, Miaowei, et al.
Published: (2026)
Flattening subtyping by eta expansion
by: Dunfield, Jana
Published: (2024)
by: Dunfield, Jana
Published: (2024)
CleverBirds: A Multiple-Choice Benchmark for Fine-grained Human Knowledge Tracing
by: Bossemeyer, Leonie, et al.
Published: (2025)
by: Bossemeyer, Leonie, et al.
Published: (2025)
Epistemic Dissonance and Modal Boundaries
by: Raileanu, Dragos
Published: (2025)
by: Raileanu, Dragos
Published: (2025)
Crowd IQ -- Aggregating Opinions to Boost Performance
by: Kosinski, Michal, et al.
Published: (2024)
by: Kosinski, Michal, et al.
Published: (2024)
DepthCues: Evaluating Monocular Depth Perception in Large Vision Models
by: Danier, Duolikun, et al.
Published: (2024)
by: Danier, Duolikun, et al.
Published: (2024)
Similar Items
-
APRES: An Agentic Paper Revision and Evaluation System
by: Zhao, Bingchen, et al.
Published: (2026) -
AI Research Agents for Machine Learning: Search, Exploration, and Generalization in MLE-bench
by: Toledo, Edan, et al.
Published: (2025) -
Bootstrapping Task Spaces for Self-Improvement
by: Jiang, Minqi, et al.
Published: (2025) -
What Does It Take to Be a Good AI Research Agent? Studying the Role of Ideation Diversity
by: Audran-Reiss, Alexis, et al.
Published: (2025) -
AIRA_2: Overcoming Bottlenecks in AI Research Agents
by: Hambardzumyan, Karen, et al.
Published: (2026)