Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks
Fuente:
arXiv
Saved in:
| Main Authors: | Qiu, Shikai, Xiao, Lechao, Wilson, Andrew Gordon, Pennington, Jeffrey, Agarwala, Atish |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
High dimensional analysis reveals conservative sharpening and a stochastic edge of stability
by: Agarwala, Atish, et al.
Published: (2024)
by: Agarwala, Atish, et al.
Published: (2024)
4+3 Phases of Compute-Optimal Neural Scaling Laws
by: Paquette, Elliot, et al.
Published: (2024)
by: Paquette, Elliot, et al.
Published: (2024)
High dimensional theory of two-phase optimizers
by: Agarwala, Atish
Published: (2026)
by: Agarwala, Atish
Published: (2026)
Per-example gradients: a new frontier for understanding and improving optimizers
by: Roulet, Vincent, et al.
Published: (2025)
by: Roulet, Vincent, et al.
Published: (2025)
To Clip or not to Clip: the Dynamics of SGD with Gradient Clipping in High-Dimensions
by: Marshall, Noah, et al.
Published: (2024)
by: Marshall, Noah, et al.
Published: (2024)
Compute Better Spent: Replacing Dense Layers with Structured Matrices
by: Qiu, Shikai, et al.
Published: (2024)
by: Qiu, Shikai, et al.
Published: (2024)
Hyperparameter Transfer Enables Consistent Gains of Matrix-Preconditioned Optimizers Across Scales
by: Qiu, Shikai, et al.
Published: (2025)
by: Qiu, Shikai, et al.
Published: (2025)
Large Language Models Are Zero-Shot Time Series Forecasters
by: Gruver, Nate, et al.
Published: (2023)
by: Gruver, Nate, et al.
Published: (2023)
Rethinking Conventional Wisdom in Machine Learning: From Generalization to Scaling
by: Xiao, Lechao
Published: (2024)
by: Xiao, Lechao
Published: (2024)
Feature learning as alignment: a structural property of gradient descent in non-linear neural networks
by: Beaglehole, Daniel, et al.
Published: (2024)
by: Beaglehole, Daniel, et al.
Published: (2024)
Exact Risk Curves of signSGD in High-Dimensions: Quantifying Preconditioning and Noise-Compression Effects
by: Xiao, Ke Liang, et al.
Published: (2024)
by: Xiao, Ke Liang, et al.
Published: (2024)
Neglected Hessian component explains mysteries in Sharpness regularization
by: Dauphin, Yann N., et al.
Published: (2024)
by: Dauphin, Yann N., et al.
Published: (2024)
From Entropy to Epiplexity: Rethinking Information for Computationally Bounded Intelligence
by: Finzi, Marc, et al.
Published: (2026)
by: Finzi, Marc, et al.
Published: (2026)
On the Interplay Between Stepsize Tuning and Progressive Sharpening
by: Roulet, Vincent, et al.
Published: (2023)
by: Roulet, Vincent, et al.
Published: (2023)
What do near-optimal learning rate schedules look like?
by: Naganuma, Hiroki, et al.
Published: (2026)
by: Naganuma, Hiroki, et al.
Published: (2026)
Customizing the Inductive Biases of Softmax Attention using Structured Matrices
by: Kuang, Yilun, et al.
Published: (2025)
by: Kuang, Yilun, et al.
Published: (2025)
Forgetting in Language Models: Capacity, Optimization, and Self-Generated Replay
by: Marek, Martin, et al.
Published: (2026)
by: Marek, Martin, et al.
Published: (2026)
Avoiding spurious sharpness minimization broadens applicability of SAM
by: Singh, Sidak Pal, et al.
Published: (2025)
by: Singh, Sidak Pal, et al.
Published: (2025)
Transferring Knowledge from Large Foundation Models to Small Downstream Models
by: Qiu, Shikai, et al.
Published: (2024)
by: Qiu, Shikai, et al.
Published: (2024)
Compute-Optimal LLMs Provably Generalize Better With Scale
by: Finzi, Marc, et al.
Published: (2025)
by: Finzi, Marc, et al.
Published: (2025)
Scaling Exponents Across Parameterizations and Optimizers
by: Everett, Katie, et al.
Published: (2024)
by: Everett, Katie, et al.
Published: (2024)
Stepping on the Edge: Curvature Aware Learning Rate Tuners
by: Roulet, Vincent, et al.
Published: (2024)
by: Roulet, Vincent, et al.
Published: (2024)
Phases of Muon: When Muon Eclipses SignSGD
by: Paquette, Elliot, et al.
Published: (2026)
by: Paquette, Elliot, et al.
Published: (2026)
Training LLMs over Neurally Compressed Text
by: Lester, Brian, et al.
Published: (2024)
by: Lester, Brian, et al.
Published: (2024)
A Study of Bayesian Neural Network Surrogates for Bayesian Optimization
by: Li, Yucen Lily, et al.
Published: (2023)
by: Li, Yucen Lily, et al.
Published: (2023)
How far away are truly hyperparameter-free learning algorithms?
by: Kasimbeg, Priya, et al.
Published: (2025)
by: Kasimbeg, Priya, et al.
Published: (2025)
Deep Learning is Not So Mysterious or Different
by: Wilson, Andrew Gordon
Published: (2025)
by: Wilson, Andrew Gordon
Published: (2025)
Scaling Laws and Compute-Optimal Training Beyond Fixed Training Durations
by: Hägele, Alexander, et al.
Published: (2024)
by: Hägele, Alexander, et al.
Published: (2024)
Out-of-Distribution Detection Methods Answer the Wrong Questions
by: Li, Yucen Lily, et al.
Published: (2025)
by: Li, Yucen Lily, et al.
Published: (2025)
Wide Neural Networks Trained with Weight Decay Provably Exhibit Neural Collapse
by: Jacot, Arthur, et al.
Published: (2024)
by: Jacot, Arthur, et al.
Published: (2024)
Training Neural Networks at Any Scale
by: Pethick, Thomas, et al.
Published: (2025)
by: Pethick, Thomas, et al.
Published: (2025)
Searching for Efficient Linear Layers over a Continuous Space of Structured Matrices
by: Potapczynski, Andres, et al.
Published: (2024)
by: Potapczynski, Andres, et al.
Published: (2024)
Neural Collapse versus Low-rank Bias: Is Deep Neural Collapse Really Optimal?
by: Súkeník, Peter, et al.
Published: (2024)
by: Súkeník, Peter, et al.
Published: (2024)
Optimal Uncertainty-guided Neural Network Training
by: Kabir, H M Dipu, et al.
Published: (2019)
by: Kabir, H M Dipu, et al.
Published: (2019)
Gradient Flow Matching for Learning Update Dynamics in Neural Network Training
by: Shou, Xiao, et al.
Published: (2025)
by: Shou, Xiao, et al.
Published: (2025)
A Training Framework for Optimal and Stable Training of Polynomial Neural Networks
by: Hossain, Forsad Al, et al.
Published: (2025)
by: Hossain, Forsad Al, et al.
Published: (2025)
Just How Flexible are Neural Networks in Practice?
by: Shwartz-Ziv, Ravid, et al.
Published: (2024)
by: Shwartz-Ziv, Ravid, et al.
Published: (2024)
Neural Network Training with Approximate Logarithmic Computations
by: Sanyal, Arnab, et al.
Published: (2019)
by: Sanyal, Arnab, et al.
Published: (2019)
On the Robustness of Neural Collapse and the Neural Collapse of Robustness
by: Su, Jingtong, et al.
Published: (2023)
by: Su, Jingtong, et al.
Published: (2023)
ReInc: Scaling Training of Dynamic Graph Neural Networks
by: Guan, Mingyu, et al.
Published: (2025)
by: Guan, Mingyu, et al.
Published: (2025)
Similar Items
-
High dimensional analysis reveals conservative sharpening and a stochastic edge of stability
by: Agarwala, Atish, et al.
Published: (2024) -
4+3 Phases of Compute-Optimal Neural Scaling Laws
by: Paquette, Elliot, et al.
Published: (2024) -
High dimensional theory of two-phase optimizers
by: Agarwala, Atish
Published: (2026) -
Per-example gradients: a new frontier for understanding and improving optimizers
by: Roulet, Vincent, et al.
Published: (2025) -
To Clip or not to Clip: the Dynamics of SGD with Gradient Clipping in High-Dimensions
by: Marshall, Noah, et al.
Published: (2024)