Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Qiu, Shikai, Xiao, Lechao, Wilson, Andrew Gordon, Pennington, Jeffrey, Agarwala, Atish |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
High dimensional analysis reveals conservative sharpening and a stochastic edge of stability
von: Agarwala, Atish, et al.
Veröffentlicht: (2024)
von: Agarwala, Atish, et al.
Veröffentlicht: (2024)
4+3 Phases of Compute-Optimal Neural Scaling Laws
von: Paquette, Elliot, et al.
Veröffentlicht: (2024)
von: Paquette, Elliot, et al.
Veröffentlicht: (2024)
High dimensional theory of two-phase optimizers
von: Agarwala, Atish
Veröffentlicht: (2026)
von: Agarwala, Atish
Veröffentlicht: (2026)
Per-example gradients: a new frontier for understanding and improving optimizers
von: Roulet, Vincent, et al.
Veröffentlicht: (2025)
von: Roulet, Vincent, et al.
Veröffentlicht: (2025)
To Clip or not to Clip: the Dynamics of SGD with Gradient Clipping in High-Dimensions
von: Marshall, Noah, et al.
Veröffentlicht: (2024)
von: Marshall, Noah, et al.
Veröffentlicht: (2024)
Compute Better Spent: Replacing Dense Layers with Structured Matrices
von: Qiu, Shikai, et al.
Veröffentlicht: (2024)
von: Qiu, Shikai, et al.
Veröffentlicht: (2024)
Hyperparameter Transfer Enables Consistent Gains of Matrix-Preconditioned Optimizers Across Scales
von: Qiu, Shikai, et al.
Veröffentlicht: (2025)
von: Qiu, Shikai, et al.
Veröffentlicht: (2025)
Large Language Models Are Zero-Shot Time Series Forecasters
von: Gruver, Nate, et al.
Veröffentlicht: (2023)
von: Gruver, Nate, et al.
Veröffentlicht: (2023)
Rethinking Conventional Wisdom in Machine Learning: From Generalization to Scaling
von: Xiao, Lechao
Veröffentlicht: (2024)
von: Xiao, Lechao
Veröffentlicht: (2024)
Feature learning as alignment: a structural property of gradient descent in non-linear neural networks
von: Beaglehole, Daniel, et al.
Veröffentlicht: (2024)
von: Beaglehole, Daniel, et al.
Veröffentlicht: (2024)
Exact Risk Curves of signSGD in High-Dimensions: Quantifying Preconditioning and Noise-Compression Effects
von: Xiao, Ke Liang, et al.
Veröffentlicht: (2024)
von: Xiao, Ke Liang, et al.
Veröffentlicht: (2024)
Neglected Hessian component explains mysteries in Sharpness regularization
von: Dauphin, Yann N., et al.
Veröffentlicht: (2024)
von: Dauphin, Yann N., et al.
Veröffentlicht: (2024)
From Entropy to Epiplexity: Rethinking Information for Computationally Bounded Intelligence
von: Finzi, Marc, et al.
Veröffentlicht: (2026)
von: Finzi, Marc, et al.
Veröffentlicht: (2026)
On the Interplay Between Stepsize Tuning and Progressive Sharpening
von: Roulet, Vincent, et al.
Veröffentlicht: (2023)
von: Roulet, Vincent, et al.
Veröffentlicht: (2023)
What do near-optimal learning rate schedules look like?
von: Naganuma, Hiroki, et al.
Veröffentlicht: (2026)
von: Naganuma, Hiroki, et al.
Veröffentlicht: (2026)
Customizing the Inductive Biases of Softmax Attention using Structured Matrices
von: Kuang, Yilun, et al.
Veröffentlicht: (2025)
von: Kuang, Yilun, et al.
Veröffentlicht: (2025)
Forgetting in Language Models: Capacity, Optimization, and Self-Generated Replay
von: Marek, Martin, et al.
Veröffentlicht: (2026)
von: Marek, Martin, et al.
Veröffentlicht: (2026)
Avoiding spurious sharpness minimization broadens applicability of SAM
von: Singh, Sidak Pal, et al.
Veröffentlicht: (2025)
von: Singh, Sidak Pal, et al.
Veröffentlicht: (2025)
Transferring Knowledge from Large Foundation Models to Small Downstream Models
von: Qiu, Shikai, et al.
Veröffentlicht: (2024)
von: Qiu, Shikai, et al.
Veröffentlicht: (2024)
Compute-Optimal LLMs Provably Generalize Better With Scale
von: Finzi, Marc, et al.
Veröffentlicht: (2025)
von: Finzi, Marc, et al.
Veröffentlicht: (2025)
Scaling Exponents Across Parameterizations and Optimizers
von: Everett, Katie, et al.
Veröffentlicht: (2024)
von: Everett, Katie, et al.
Veröffentlicht: (2024)
Stepping on the Edge: Curvature Aware Learning Rate Tuners
von: Roulet, Vincent, et al.
Veröffentlicht: (2024)
von: Roulet, Vincent, et al.
Veröffentlicht: (2024)
Phases of Muon: When Muon Eclipses SignSGD
von: Paquette, Elliot, et al.
Veröffentlicht: (2026)
von: Paquette, Elliot, et al.
Veröffentlicht: (2026)
Training LLMs over Neurally Compressed Text
von: Lester, Brian, et al.
Veröffentlicht: (2024)
von: Lester, Brian, et al.
Veröffentlicht: (2024)
A Study of Bayesian Neural Network Surrogates for Bayesian Optimization
von: Li, Yucen Lily, et al.
Veröffentlicht: (2023)
von: Li, Yucen Lily, et al.
Veröffentlicht: (2023)
How far away are truly hyperparameter-free learning algorithms?
von: Kasimbeg, Priya, et al.
Veröffentlicht: (2025)
von: Kasimbeg, Priya, et al.
Veröffentlicht: (2025)
Deep Learning is Not So Mysterious or Different
von: Wilson, Andrew Gordon
Veröffentlicht: (2025)
von: Wilson, Andrew Gordon
Veröffentlicht: (2025)
Scaling Laws and Compute-Optimal Training Beyond Fixed Training Durations
von: Hägele, Alexander, et al.
Veröffentlicht: (2024)
von: Hägele, Alexander, et al.
Veröffentlicht: (2024)
Out-of-Distribution Detection Methods Answer the Wrong Questions
von: Li, Yucen Lily, et al.
Veröffentlicht: (2025)
von: Li, Yucen Lily, et al.
Veröffentlicht: (2025)
Wide Neural Networks Trained with Weight Decay Provably Exhibit Neural Collapse
von: Jacot, Arthur, et al.
Veröffentlicht: (2024)
von: Jacot, Arthur, et al.
Veröffentlicht: (2024)
Training Neural Networks at Any Scale
von: Pethick, Thomas, et al.
Veröffentlicht: (2025)
von: Pethick, Thomas, et al.
Veröffentlicht: (2025)
Searching for Efficient Linear Layers over a Continuous Space of Structured Matrices
von: Potapczynski, Andres, et al.
Veröffentlicht: (2024)
von: Potapczynski, Andres, et al.
Veröffentlicht: (2024)
Neural Collapse versus Low-rank Bias: Is Deep Neural Collapse Really Optimal?
von: Súkeník, Peter, et al.
Veröffentlicht: (2024)
von: Súkeník, Peter, et al.
Veröffentlicht: (2024)
Optimal Uncertainty-guided Neural Network Training
von: Kabir, H M Dipu, et al.
Veröffentlicht: (2019)
von: Kabir, H M Dipu, et al.
Veröffentlicht: (2019)
Gradient Flow Matching for Learning Update Dynamics in Neural Network Training
von: Shou, Xiao, et al.
Veröffentlicht: (2025)
von: Shou, Xiao, et al.
Veröffentlicht: (2025)
A Training Framework for Optimal and Stable Training of Polynomial Neural Networks
von: Hossain, Forsad Al, et al.
Veröffentlicht: (2025)
von: Hossain, Forsad Al, et al.
Veröffentlicht: (2025)
Just How Flexible are Neural Networks in Practice?
von: Shwartz-Ziv, Ravid, et al.
Veröffentlicht: (2024)
von: Shwartz-Ziv, Ravid, et al.
Veröffentlicht: (2024)
Neural Network Training with Approximate Logarithmic Computations
von: Sanyal, Arnab, et al.
Veröffentlicht: (2019)
von: Sanyal, Arnab, et al.
Veröffentlicht: (2019)
On the Robustness of Neural Collapse and the Neural Collapse of Robustness
von: Su, Jingtong, et al.
Veröffentlicht: (2023)
von: Su, Jingtong, et al.
Veröffentlicht: (2023)
ReInc: Scaling Training of Dynamic Graph Neural Networks
von: Guan, Mingyu, et al.
Veröffentlicht: (2025)
von: Guan, Mingyu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
High dimensional analysis reveals conservative sharpening and a stochastic edge of stability
von: Agarwala, Atish, et al.
Veröffentlicht: (2024) -
4+3 Phases of Compute-Optimal Neural Scaling Laws
von: Paquette, Elliot, et al.
Veröffentlicht: (2024) -
High dimensional theory of two-phase optimizers
von: Agarwala, Atish
Veröffentlicht: (2026) -
Per-example gradients: a new frontier for understanding and improving optimizers
von: Roulet, Vincent, et al.
Veröffentlicht: (2025) -
To Clip or not to Clip: the Dynamics of SGD with Gradient Clipping in High-Dimensions
von: Marshall, Noah, et al.
Veröffentlicht: (2024)