Scaling with Collapse: Efficient and Predictable Training of LLM Families

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bergsma, Shane, Zhang, Bin Claire, Dey, Nolan, Muhammad, Shaheer, Gosal, Gurpreet, Hestness, Joel
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915826473369600
author Bergsma, Shane
Zhang, Bin Claire
Dey, Nolan
Muhammad, Shaheer
Gosal, Gurpreet
Hestness, Joel
author_facet Bergsma, Shane
Zhang, Bin Claire
Dey, Nolan
Muhammad, Shaheer
Gosal, Gurpreet
Hestness, Joel
contents Effective LLM training depends on predictable scaling of key quantities -- such as final loss and optimal hyperparameters -- with model and dataset size. Qiu et al. (2025) recently showed that this predictability can extend beyond scalars: whole training loss curves can *collapse* onto a universal trajectory after a simple normalization. What remains unclear is whether this phenomenon persists for LLM families trained under *practical scaling recipes*, where width, depth, learning rate, batch size, and weight decay are scaled jointly. We show that it does: loss curves collapse across scales precisely when optimization hyperparameters are set optimally for the given data budget, in accordance with recent empirical scaling laws. Collapse therefore emerges as a signature of compute-efficient training. We demonstrate two applications at scale: (1) deviation-from-collapse provides a sensitive, early diagnostic of training pathologies, and (2) predictability of collapsed curves enables early stopping in large-scale hyperparameter tuning. Finally, we train a competitive LLM family, *Celerity*, using these insights, establishing collapse as an effective tool for developing efficient LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2509_25087
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Scaling with Collapse: Efficient and Predictable Training of LLM Families
Bergsma, Shane
Zhang, Bin Claire
Dey, Nolan
Muhammad, Shaheer
Gosal, Gurpreet
Hestness, Joel
Machine Learning
Artificial Intelligence
Computation and Language
Effective LLM training depends on predictable scaling of key quantities -- such as final loss and optimal hyperparameters -- with model and dataset size. Qiu et al. (2025) recently showed that this predictability can extend beyond scalars: whole training loss curves can *collapse* onto a universal trajectory after a simple normalization. What remains unclear is whether this phenomenon persists for LLM families trained under *practical scaling recipes*, where width, depth, learning rate, batch size, and weight decay are scaled jointly. We show that it does: loss curves collapse across scales precisely when optimization hyperparameters are set optimally for the given data budget, in accordance with recent empirical scaling laws. Collapse therefore emerges as a signature of compute-efficient training. We demonstrate two applications at scale: (1) deviation-from-collapse provides a sensitive, early diagnostic of training pathologies, and (2) predictability of collapsed curves enables early stopping in large-scale hyperparameter tuning. Finally, we train a competitive LLM family, *Celerity*, using these insights, establishing collapse as an effective tool for developing efficient LLMs.
title Scaling with Collapse: Efficient and Predictable Training of LLM Families
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2509.25087