When is Warmstarting Effective for Scaling Language Models?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mallik, Neeratyoy, Janowski, Maciej, Hog, Johannes, Rakotoarison, Herilalaina, Grabocka, Josif, Hutter, Frank, Klein, Aaron
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910216179679232
author Mallik, Neeratyoy
Janowski, Maciej
Hog, Johannes
Rakotoarison, Herilalaina
Grabocka, Josif
Hutter, Frank
Klein, Aaron
author_facet Mallik, Neeratyoy
Janowski, Maciej
Hog, Johannes
Rakotoarison, Herilalaina
Grabocka, Josif
Hutter, Frank
Klein, Aaron
contents Model growth from a given checkpoint aims to accelerate training of a larger model, offering potential resource savings. Despite recent interest, warmstarting has seen limited practical adoption in large-scale training. We attribute this to two underexplored factors: (1) an overemphasis on preserving the smaller model's performance at initialization, which constrains operator design for new architectures, and (2) insufficient analysis of how growth interacts with hyperparameters and scaling behavior, compounded by inconsistent growth factors across the literature. We show that preserving the base model's initial post-growth performance is not necessary for strong final performance, and that simple, architecture-agnostic growth strategies can outperform more complex warmstarting operators. Crucially, we empirically identify an upper bound on the growth factor $g$ beyond which training from scratch is more efficient. We observe this across multiple ablation setups. Notably, this limit is also present, but unreported, in prior published results. Across our experiments on dense MLPs and dense language models, we find that a $2\times$ growth factor is the most reliable in yielding convergence speedups, with gains most pronounced under 20 tokens/parameter budgets and diminishing as budget increases. We fit scaling laws over these observations to provide predictive guidance for practitioners deciding when and how much to grow. Together, our analysis provides practical guidelines and empirical limits for model growth.
format Preprint
id arxiv_https___arxiv_org_abs_2605_13405
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle When is Warmstarting Effective for Scaling Language Models?
Mallik, Neeratyoy
Janowski, Maciej
Hog, Johannes
Rakotoarison, Herilalaina
Grabocka, Josif
Hutter, Frank
Klein, Aaron
Machine Learning
Model growth from a given checkpoint aims to accelerate training of a larger model, offering potential resource savings. Despite recent interest, warmstarting has seen limited practical adoption in large-scale training. We attribute this to two underexplored factors: (1) an overemphasis on preserving the smaller model's performance at initialization, which constrains operator design for new architectures, and (2) insufficient analysis of how growth interacts with hyperparameters and scaling behavior, compounded by inconsistent growth factors across the literature. We show that preserving the base model's initial post-growth performance is not necessary for strong final performance, and that simple, architecture-agnostic growth strategies can outperform more complex warmstarting operators. Crucially, we empirically identify an upper bound on the growth factor $g$ beyond which training from scratch is more efficient. We observe this across multiple ablation setups. Notably, this limit is also present, but unreported, in prior published results. Across our experiments on dense MLPs and dense language models, we find that a $2\times$ growth factor is the most reliable in yielding convergence speedups, with gains most pronounced under 20 tokens/parameter budgets and diminishing as budget increases. We fit scaling laws over these observations to provide predictive guidance for practitioners deciding when and how much to grow. Together, our analysis provides practical guidelines and empirical limits for model growth.
title When is Warmstarting Effective for Scaling Language Models?
topic Machine Learning
url https://arxiv.org/abs/2605.13405