Test-Time Scaling Makes Overtraining Compute-Optimal

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Roberts, Nicholas, Cho, Sungjun, Gao, Zhiqi, Huang, Tzu-Heng, Wu, Albert, Orlanski, Gabriel, Trost, Avi, Buchanan, Kelly, Albarghouthi, Aws, Sala, Frederic
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914439222001664
author Roberts, Nicholas
Cho, Sungjun
Gao, Zhiqi
Huang, Tzu-Heng
Wu, Albert
Orlanski, Gabriel
Trost, Avi
Buchanan, Kelly
Albarghouthi, Aws
Sala, Frederic
author_facet Roberts, Nicholas
Cho, Sungjun
Gao, Zhiqi
Huang, Tzu-Heng
Wu, Albert
Orlanski, Gabriel
Trost, Avi
Buchanan, Kelly
Albarghouthi, Aws
Sala, Frederic
contents Modern LLMs scale at test-time, e.g. via repeated sampling, where inference cost grows with model size and the number of samples. This creates a trade-off that pretraining scaling laws, such as Chinchilla, do not address. We present Train-to-Test ($T^2$) scaling laws that jointly optimize model size, training tokens, and number of inference samples under fixed end-to-end budgets. $T^2$ modernizes pretraining scaling laws with pass@$k$ modeling used for test-time scaling, then jointly optimizes pretraining and test-time decisions. Forecasts from $T^2$ are robust over distinct modeling approaches: measuring joint scaling effect on the task loss and modeling impact on task accuracy. Across eight downstream tasks, we find that when accounting for inference cost, optimal pretraining decisions shift radically into the overtraining regime, well-outside of the range of standard pretraining scaling suites. We validate our results by pretraining heavily overtrained models in the optimal region that $T^2$ scaling forecasts, confirming their substantially stronger performance compared to pretraining scaling alone. Finally, as frontier LLMs are post-trained, we show that our findings survive the post-training stage, making $T^2$ scaling meaningful in modern deployments.
format Preprint
id arxiv_https___arxiv_org_abs_2604_01411
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Test-Time Scaling Makes Overtraining Compute-Optimal
Roberts, Nicholas
Cho, Sungjun
Gao, Zhiqi
Huang, Tzu-Heng
Wu, Albert
Orlanski, Gabriel
Trost, Avi
Buchanan, Kelly
Albarghouthi, Aws
Sala, Frederic
Machine Learning
Computation and Language
Modern LLMs scale at test-time, e.g. via repeated sampling, where inference cost grows with model size and the number of samples. This creates a trade-off that pretraining scaling laws, such as Chinchilla, do not address. We present Train-to-Test ($T^2$) scaling laws that jointly optimize model size, training tokens, and number of inference samples under fixed end-to-end budgets. $T^2$ modernizes pretraining scaling laws with pass@$k$ modeling used for test-time scaling, then jointly optimizes pretraining and test-time decisions. Forecasts from $T^2$ are robust over distinct modeling approaches: measuring joint scaling effect on the task loss and modeling impact on task accuracy. Across eight downstream tasks, we find that when accounting for inference cost, optimal pretraining decisions shift radically into the overtraining regime, well-outside of the range of standard pretraining scaling suites. We validate our results by pretraining heavily overtrained models in the optimal region that $T^2$ scaling forecasts, confirming their substantially stronger performance compared to pretraining scaling alone. Finally, as frontier LLMs are post-trained, we show that our findings survive the post-training stage, making $T^2$ scaling meaningful in modern deployments.
title Test-Time Scaling Makes Overtraining Compute-Optimal
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2604.01411