A Dynamical Model of Neural Scaling Laws
Fuente:
arXiv
Guardado en:
| Autores principales: | Bordelon, Blake, Atanasov, Alexander, Pehlevan, Cengiz |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
How Feature Learning Can Improve Neural Scaling Laws
por: Bordelon, Blake, et al.
Publicado: (2024)
por: Bordelon, Blake, et al.
Publicado: (2024)
Disordered Dynamics in High Dimensions: Connections to Random Matrices and Machine Learning
por: Bordelon, Blake, et al.
Publicado: (2026)
por: Bordelon, Blake, et al.
Publicado: (2026)
Deep Linear Network Training Dynamics from Random Initialization: Data, Width, Depth, and Hyperparameter Transfer
por: Bordelon, Blake, et al.
Publicado: (2025)
por: Bordelon, Blake, et al.
Publicado: (2025)
Theory of Scaling Laws for In-Context Regression: Depth, Width, Context and Time
por: Bordelon, Blake, et al.
Publicado: (2025)
por: Bordelon, Blake, et al.
Publicado: (2025)
Two-Point Deterministic Equivalence for Stochastic Gradient Dynamics in Linear Models
por: Atanasov, Alexander, et al.
Publicado: (2025)
por: Atanasov, Alexander, et al.
Publicado: (2025)
Infinite Limits of Multi-head Transformer Dynamics
por: Bordelon, Blake, et al.
Publicado: (2024)
por: Bordelon, Blake, et al.
Publicado: (2024)
Adaptive kernel predictors from feature-learning infinite limits of neural networks
por: Lauditi, Clarissa, et al.
Publicado: (2025)
por: Lauditi, Clarissa, et al.
Publicado: (2025)
Transfer Learning in Infinite Width Feature Learning Networks
por: Lauditi, Clarissa, et al.
Publicado: (2025)
por: Lauditi, Clarissa, et al.
Publicado: (2025)
Spectral Dynamics in Deep Networks: Feature Learning, Outlier Escape, and Learning Rate Transfer
por: Lauditi, Clarissa, et al.
Publicado: (2026)
por: Lauditi, Clarissa, et al.
Publicado: (2026)
Grokking as the Transition from Lazy to Rich Training Dynamics
por: Kumar, Tanishq, et al.
Publicado: (2023)
por: Kumar, Tanishq, et al.
Publicado: (2023)
Scaling and renormalization in high-dimensional regression
por: Atanasov, Alexander, et al.
Publicado: (2024)
por: Atanasov, Alexander, et al.
Publicado: (2024)
Dynamically Learning to Integrate in Recurrent Neural Networks
por: Bordelon, Blake, et al.
Publicado: (2025)
por: Bordelon, Blake, et al.
Publicado: (2025)
Theory of Optimal Learning Rate Schedules and Scaling Laws for a Random Feature Model
por: Bordelon, Blake, et al.
Publicado: (2026)
por: Bordelon, Blake, et al.
Publicado: (2026)
Risk and cross validation in ridge regression with correlated samples
por: Atanasov, Alexander, et al.
Publicado: (2024)
por: Atanasov, Alexander, et al.
Publicado: (2024)
No Free Lunch From Random Feature Ensembles: Scaling Laws and Near-Optimality Conditions
por: Ruben, Benjamin S., et al.
Publicado: (2024)
por: Ruben, Benjamin S., et al.
Publicado: (2024)
Nadaraya-Watson kernel smoothing as a random energy model
por: Zavatone-Veth, Jacob A., et al.
Publicado: (2024)
por: Zavatone-Veth, Jacob A., et al.
Publicado: (2024)
A solvable model of learning generative diffusion: theory and insights
por: Cui, Hugo, et al.
Publicado: (2025)
por: Cui, Hugo, et al.
Publicado: (2025)
Learning Curves for Noisy Heterogeneous Feature-Subsampled Ridge Ensembles
por: Ruben, Benjamin S., et al.
Publicado: (2023)
por: Ruben, Benjamin S., et al.
Publicado: (2023)
How does training shape the Riemannian geometry of neural network representations?
por: Zavatone-Veth, Jacob A., et al.
Publicado: (2023)
por: Zavatone-Veth, Jacob A., et al.
Publicado: (2023)
Asymptotic theory of in-context learning by linear attention
por: Lu, Yue M., et al.
Publicado: (2024)
por: Lu, Yue M., et al.
Publicado: (2024)
Explaining Neural Scaling Laws
por: Bahri, Yasaman, et al.
Publicado: (2021)
por: Bahri, Yasaman, et al.
Publicado: (2021)
A note on the dynamics of extended-context disordered kinetic spin models
por: Zavatone-Veth, Jacob A., et al.
Publicado: (2025)
por: Zavatone-Veth, Jacob A., et al.
Publicado: (2025)
Neural Scaling Laws Rooted in the Data Distribution
por: Brill, Ari
Publicado: (2024)
por: Brill, Ari
Publicado: (2024)
The Quantization Model of Neural Scaling
por: Michaud, Eric J., et al.
Publicado: (2023)
por: Michaud, Eric J., et al.
Publicado: (2023)
Scaling Laws and Representation Learning in Simple Hierarchical Languages: Transformers vs. Convolutional Architectures
por: Cagnetta, Francesco, et al.
Publicado: (2025)
por: Cagnetta, Francesco, et al.
Publicado: (2025)
Dropout Universality: Scaling Laws and Optimal Scheduling at the Edge-of-Chaos
por: Sarmiento, Lucas Fernandez
Publicado: (2026)
por: Sarmiento, Lucas Fernandez
Publicado: (2026)
Scaling Laws and Spectra of Shallow Neural Networks in the Feature Learning Regime
por: Defilippis, Leonardo, et al.
Publicado: (2025)
por: Defilippis, Leonardo, et al.
Publicado: (2025)
Universal Scaling Laws of Absorbing Phase Transitions in Artificial Deep Neural Networks
por: Tamai, Keiichi, et al.
Publicado: (2023)
por: Tamai, Keiichi, et al.
Publicado: (2023)
Dynamical Mean-Field Theory of Self-Attention Neural Networks
por: Poc-López, Ángel, et al.
Publicado: (2024)
por: Poc-López, Ángel, et al.
Publicado: (2024)
Growing Neural Networks: Dynamic Evolution through Gradient Descent
por: Radhakrishnan, Anil, et al.
Publicado: (2025)
por: Radhakrishnan, Anil, et al.
Publicado: (2025)
Controlled Langevin Dynamics for Sampling of Feedforward Neural Networks Trained with Minibatches
por: Zambon, Alessandro, et al.
Publicado: (2026)
por: Zambon, Alessandro, et al.
Publicado: (2026)
Dynamical Regimes of Multimodal Diffusion Models
por: Albrychiewicz, Emil, et al.
Publicado: (2026)
por: Albrychiewicz, Emil, et al.
Publicado: (2026)
Asymmetric Scaling Laws from Sparse Features
por: Sous, John, et al.
Publicado: (2026)
por: Sous, John, et al.
Publicado: (2026)
The Rules-and-Facts Model for Simultaneous Generalization and Memorization in Neural Networks
por: Farné, Gabriele, et al.
Publicado: (2026)
por: Farné, Gabriele, et al.
Publicado: (2026)
The Interplay of Data Structure and Imbalance in the Learning Dynamics of Diffusion Models
por: Nicoletti, Flavio, et al.
Publicado: (2026)
por: Nicoletti, Flavio, et al.
Publicado: (2026)
Training Dynamics of Nonlinear Contrastive Learning Model in the High Dimensional Limit
por: Meng, Lineghuan, et al.
Publicado: (2024)
por: Meng, Lineghuan, et al.
Publicado: (2024)
Siamese Neural Network for Label-Efficient Critical Phenomena Prediction in 3D Percolation Models
por: Wang, Shanshan, et al.
Publicado: (2025)
por: Wang, Shanshan, et al.
Publicado: (2025)
Why Diffusion Models Don't Memorize: The Role of Implicit Dynamical Regularization in Training
por: Bonnaire, Tony, et al.
Publicado: (2025)
por: Bonnaire, Tony, et al.
Publicado: (2025)
Precise Dynamics of Diagonal Linear Networks: A Unifying Analysis by Dynamical Mean-Field Theory
por: Nishiyama, Sota, et al.
Publicado: (2025)
por: Nishiyama, Sota, et al.
Publicado: (2025)
A Boundary-Layer Mechanism for One-Third Scaling in Online Softmax Classification
por: Kühn, Marcel, et al.
Publicado: (2026)
por: Kühn, Marcel, et al.
Publicado: (2026)
Ejemplares similares
-
How Feature Learning Can Improve Neural Scaling Laws
por: Bordelon, Blake, et al.
Publicado: (2024) -
Disordered Dynamics in High Dimensions: Connections to Random Matrices and Machine Learning
por: Bordelon, Blake, et al.
Publicado: (2026) -
Deep Linear Network Training Dynamics from Random Initialization: Data, Width, Depth, and Hyperparameter Transfer
por: Bordelon, Blake, et al.
Publicado: (2025) -
Theory of Scaling Laws for In-Context Regression: Depth, Width, Context and Time
por: Bordelon, Blake, et al.
Publicado: (2025) -
Two-Point Deterministic Equivalence for Stochastic Gradient Dynamics in Linear Models
por: Atanasov, Alexander, et al.
Publicado: (2025)