How Feature Learning Can Improve Neural Scaling Laws
Fuente:
arXiv
Saved in:
| Main Authors: | Bordelon, Blake, Atanasov, Alexander, Pehlevan, Cengiz |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Dynamical Model of Neural Scaling Laws
by: Bordelon, Blake, et al.
Published: (2024)
by: Bordelon, Blake, et al.
Published: (2024)
Disordered Dynamics in High Dimensions: Connections to Random Matrices and Machine Learning
by: Bordelon, Blake, et al.
Published: (2026)
by: Bordelon, Blake, et al.
Published: (2026)
Transfer Learning in Infinite Width Feature Learning Networks
by: Lauditi, Clarissa, et al.
Published: (2025)
by: Lauditi, Clarissa, et al.
Published: (2025)
Theory of Scaling Laws for In-Context Regression: Depth, Width, Context and Time
by: Bordelon, Blake, et al.
Published: (2025)
by: Bordelon, Blake, et al.
Published: (2025)
Deep Linear Network Training Dynamics from Random Initialization: Data, Width, Depth, and Hyperparameter Transfer
by: Bordelon, Blake, et al.
Published: (2025)
by: Bordelon, Blake, et al.
Published: (2025)
Spectral Dynamics in Deep Networks: Feature Learning, Outlier Escape, and Learning Rate Transfer
by: Lauditi, Clarissa, et al.
Published: (2026)
by: Lauditi, Clarissa, et al.
Published: (2026)
Adaptive kernel predictors from feature-learning infinite limits of neural networks
by: Lauditi, Clarissa, et al.
Published: (2025)
by: Lauditi, Clarissa, et al.
Published: (2025)
Two-Point Deterministic Equivalence for Stochastic Gradient Dynamics in Linear Models
by: Atanasov, Alexander, et al.
Published: (2025)
by: Atanasov, Alexander, et al.
Published: (2025)
Infinite Limits of Multi-head Transformer Dynamics
by: Bordelon, Blake, et al.
Published: (2024)
by: Bordelon, Blake, et al.
Published: (2024)
Scaling and renormalization in high-dimensional regression
by: Atanasov, Alexander, et al.
Published: (2024)
by: Atanasov, Alexander, et al.
Published: (2024)
Theory of Optimal Learning Rate Schedules and Scaling Laws for a Random Feature Model
by: Bordelon, Blake, et al.
Published: (2026)
by: Bordelon, Blake, et al.
Published: (2026)
Grokking as the Transition from Lazy to Rich Training Dynamics
by: Kumar, Tanishq, et al.
Published: (2023)
by: Kumar, Tanishq, et al.
Published: (2023)
Dynamically Learning to Integrate in Recurrent Neural Networks
by: Bordelon, Blake, et al.
Published: (2025)
by: Bordelon, Blake, et al.
Published: (2025)
Risk and cross validation in ridge regression with correlated samples
by: Atanasov, Alexander, et al.
Published: (2024)
by: Atanasov, Alexander, et al.
Published: (2024)
No Free Lunch From Random Feature Ensembles: Scaling Laws and Near-Optimality Conditions
by: Ruben, Benjamin S., et al.
Published: (2024)
by: Ruben, Benjamin S., et al.
Published: (2024)
Learning Curves for Noisy Heterogeneous Feature-Subsampled Ridge Ensembles
by: Ruben, Benjamin S., et al.
Published: (2023)
by: Ruben, Benjamin S., et al.
Published: (2023)
Nadaraya-Watson kernel smoothing as a random energy model
by: Zavatone-Veth, Jacob A., et al.
Published: (2024)
by: Zavatone-Veth, Jacob A., et al.
Published: (2024)
A solvable model of learning generative diffusion: theory and insights
by: Cui, Hugo, et al.
Published: (2025)
by: Cui, Hugo, et al.
Published: (2025)
How does training shape the Riemannian geometry of neural network representations?
by: Zavatone-Veth, Jacob A., et al.
Published: (2023)
by: Zavatone-Veth, Jacob A., et al.
Published: (2023)
Asymptotic theory of in-context learning by linear attention
by: Lu, Yue M., et al.
Published: (2024)
by: Lu, Yue M., et al.
Published: (2024)
Explaining Neural Scaling Laws
by: Bahri, Yasaman, et al.
Published: (2021)
by: Bahri, Yasaman, et al.
Published: (2021)
Neural Scaling Laws Rooted in the Data Distribution
by: Brill, Ari
Published: (2024)
by: Brill, Ari
Published: (2024)
A note on the dynamics of extended-context disordered kinetic spin models
by: Zavatone-Veth, Jacob A., et al.
Published: (2025)
by: Zavatone-Veth, Jacob A., et al.
Published: (2025)
Scaling Laws and Spectra of Shallow Neural Networks in the Feature Learning Regime
by: Defilippis, Leonardo, et al.
Published: (2025)
by: Defilippis, Leonardo, et al.
Published: (2025)
From Kernels to Features: A Multi-Scale Adaptive Theory of Feature Learning
by: Rubin, Noa, et al.
Published: (2025)
by: Rubin, Noa, et al.
Published: (2025)
Asymmetric Scaling Laws from Sparse Features
by: Sous, John, et al.
Published: (2026)
by: Sous, John, et al.
Published: (2026)
Scaling Laws and Representation Learning in Simple Hierarchical Languages: Transformers vs. Convolutional Architectures
by: Cagnetta, Francesco, et al.
Published: (2025)
by: Cagnetta, Francesco, et al.
Published: (2025)
Deep Learning as Neural Low-Degree Filtering: A Spectral Theory of Hierarchical Feature Learning
by: Dandi, Yatin, et al.
Published: (2026)
by: Dandi, Yatin, et al.
Published: (2026)
The Quantization Model of Neural Scaling
by: Michaud, Eric J., et al.
Published: (2023)
by: Michaud, Eric J., et al.
Published: (2023)
Statistical Physics of Deep Neural Networks: Generalization Capability, Beyond the Infinite Width, and Feature Learning
by: Ariosto, Sebastiano
Published: (2025)
by: Ariosto, Sebastiano
Published: (2025)
Dropout Universality: Scaling Laws and Optimal Scheduling at the Edge-of-Chaos
by: Sarmiento, Lucas Fernandez
Published: (2026)
by: Sarmiento, Lucas Fernandez
Published: (2026)
Universal Scaling Laws of Absorbing Phase Transitions in Artificial Deep Neural Networks
by: Tamai, Keiichi, et al.
Published: (2023)
by: Tamai, Keiichi, et al.
Published: (2023)
How Deep Networks Learn Sparse and Hierarchical Data: the Sparse Random Hierarchy Model
by: Tomasini, Umberto, et al.
Published: (2024)
by: Tomasini, Umberto, et al.
Published: (2024)
Analog Physical Systems Can Exhibit Double Descent
by: Dillavou, Sam, et al.
Published: (2025)
by: Dillavou, Sam, et al.
Published: (2025)
When Can You Get Away with Low Memory Adam?
by: Kalra, Dayal Singh, et al.
Published: (2025)
by: Kalra, Dayal Singh, et al.
Published: (2025)
(How) Can Transformers Predict Pseudo-Random Numbers?
by: Tao, Tao, et al.
Published: (2025)
by: Tao, Tao, et al.
Published: (2025)
Formation of Representations in Neural Networks
by: Ziyin, Liu, et al.
Published: (2024)
by: Ziyin, Liu, et al.
Published: (2024)
Stochastic Gradient Flow Dynamics of Test Risk and its Exact Solution for Weak Features
by: Veiga, Rodrigo, et al.
Published: (2024)
by: Veiga, Rodrigo, et al.
Published: (2024)
Initial Guessing Bias: How Untrained Networks Favor Some Classes
by: Francazi, Emanuele, et al.
Published: (2023)
by: Francazi, Emanuele, et al.
Published: (2023)
Grokking vs. Learning: Same Features, Different Encodings
by: Manning-Coe, Dmitry, et al.
Published: (2025)
by: Manning-Coe, Dmitry, et al.
Published: (2025)
Similar Items
-
A Dynamical Model of Neural Scaling Laws
by: Bordelon, Blake, et al.
Published: (2024) -
Disordered Dynamics in High Dimensions: Connections to Random Matrices and Machine Learning
by: Bordelon, Blake, et al.
Published: (2026) -
Transfer Learning in Infinite Width Feature Learning Networks
by: Lauditi, Clarissa, et al.
Published: (2025) -
Theory of Scaling Laws for In-Context Regression: Depth, Width, Context and Time
by: Bordelon, Blake, et al.
Published: (2025) -
Deep Linear Network Training Dynamics from Random Initialization: Data, Width, Depth, and Hyperparameter Transfer
by: Bordelon, Blake, et al.
Published: (2025)