Theory of Optimal Learning Rate Schedules and Scaling Laws for a Random Feature Model
Fuente:
arXiv
Guardado en:
| Autores principales: | Bordelon, Blake, Mori, Francesco |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
How Feature Learning Can Improve Neural Scaling Laws
por: Bordelon, Blake, et al.
Publicado: (2024)
por: Bordelon, Blake, et al.
Publicado: (2024)
A Dynamical Model of Neural Scaling Laws
por: Bordelon, Blake, et al.
Publicado: (2024)
por: Bordelon, Blake, et al.
Publicado: (2024)
Theory of Scaling Laws for In-Context Regression: Depth, Width, Context and Time
por: Bordelon, Blake, et al.
Publicado: (2025)
por: Bordelon, Blake, et al.
Publicado: (2025)
Disordered Dynamics in High Dimensions: Connections to Random Matrices and Machine Learning
por: Bordelon, Blake, et al.
Publicado: (2026)
por: Bordelon, Blake, et al.
Publicado: (2026)
Transfer Learning in Infinite Width Feature Learning Networks
por: Lauditi, Clarissa, et al.
Publicado: (2025)
por: Lauditi, Clarissa, et al.
Publicado: (2025)
Deep Linear Network Training Dynamics from Random Initialization: Data, Width, Depth, and Hyperparameter Transfer
por: Bordelon, Blake, et al.
Publicado: (2025)
por: Bordelon, Blake, et al.
Publicado: (2025)
Spectral Dynamics in Deep Networks: Feature Learning, Outlier Escape, and Learning Rate Transfer
por: Lauditi, Clarissa, et al.
Publicado: (2026)
por: Lauditi, Clarissa, et al.
Publicado: (2026)
Adaptive kernel predictors from feature-learning infinite limits of neural networks
por: Lauditi, Clarissa, et al.
Publicado: (2025)
por: Lauditi, Clarissa, et al.
Publicado: (2025)
No Free Lunch From Random Feature Ensembles: Scaling Laws and Near-Optimality Conditions
por: Ruben, Benjamin S., et al.
Publicado: (2024)
por: Ruben, Benjamin S., et al.
Publicado: (2024)
Infinite Limits of Multi-head Transformer Dynamics
por: Bordelon, Blake, et al.
Publicado: (2024)
por: Bordelon, Blake, et al.
Publicado: (2024)
Dropout Universality: Scaling Laws and Optimal Scheduling at the Edge-of-Chaos
por: Sarmiento, Lucas Fernandez
Publicado: (2026)
por: Sarmiento, Lucas Fernandez
Publicado: (2026)
Grokking as the Transition from Lazy to Rich Training Dynamics
por: Kumar, Tanishq, et al.
Publicado: (2023)
por: Kumar, Tanishq, et al.
Publicado: (2023)
Two-Point Deterministic Equivalence for Stochastic Gradient Dynamics in Linear Models
por: Atanasov, Alexander, et al.
Publicado: (2025)
por: Atanasov, Alexander, et al.
Publicado: (2025)
Optimal Protocols for Continual Learning via Statistical Physics and Control Theory
por: Mori, Francesco, et al.
Publicado: (2024)
por: Mori, Francesco, et al.
Publicado: (2024)
Dynamically Learning to Integrate in Recurrent Neural Networks
por: Bordelon, Blake, et al.
Publicado: (2025)
por: Bordelon, Blake, et al.
Publicado: (2025)
From Kernels to Features: A Multi-Scale Adaptive Theory of Feature Learning
por: Rubin, Noa, et al.
Publicado: (2025)
por: Rubin, Noa, et al.
Publicado: (2025)
Scaling Laws and Representation Learning in Simple Hierarchical Languages: Transformers vs. Convolutional Architectures
por: Cagnetta, Francesco, et al.
Publicado: (2025)
por: Cagnetta, Francesco, et al.
Publicado: (2025)
Explaining Neural Scaling Laws
por: Bahri, Yasaman, et al.
Publicado: (2021)
por: Bahri, Yasaman, et al.
Publicado: (2021)
Scaling Laws and Spectra of Shallow Neural Networks in the Feature Learning Regime
por: Defilippis, Leonardo, et al.
Publicado: (2025)
por: Defilippis, Leonardo, et al.
Publicado: (2025)
Asymmetric Scaling Laws from Sparse Features
por: Sous, John, et al.
Publicado: (2026)
por: Sous, John, et al.
Publicado: (2026)
Neural Scaling Laws Rooted in the Data Distribution
por: Brill, Ari
Publicado: (2024)
por: Brill, Ari
Publicado: (2024)
Deep Learning as Neural Low-Degree Filtering: A Spectral Theory of Hierarchical Feature Learning
por: Dandi, Yatin, et al.
Publicado: (2026)
por: Dandi, Yatin, et al.
Publicado: (2026)
A Random-Matrix Criterion for Initializing Gated Recurrent Neural Networks
por: Fioratti, Tommaso, et al.
Publicado: (2026)
por: Fioratti, Tommaso, et al.
Publicado: (2026)
Modeling Structured Data Learning with Restricted Boltzmann Machines in the Teacher-Student Setting
por: Thériault, Robin, et al.
Publicado: (2024)
por: Thériault, Robin, et al.
Publicado: (2024)
How Deep Networks Learn Sparse and Hierarchical Data: the Sparse Random Hierarchy Model
por: Tomasini, Umberto, et al.
Publicado: (2024)
por: Tomasini, Umberto, et al.
Publicado: (2024)
Why Warmup the Learning Rate? Underlying Mechanisms and Improvements
por: Kalra, Dayal Singh, et al.
Publicado: (2024)
por: Kalra, Dayal Singh, et al.
Publicado: (2024)
Asymptotics of Learning with Deep Structured (Random) Features
por: Schröder, Dominik, et al.
Publicado: (2024)
por: Schröder, Dominik, et al.
Publicado: (2024)
Analytic theory of dropout regularization
por: Mori, Francesco, et al.
Publicado: (2025)
por: Mori, Francesco, et al.
Publicado: (2025)
Random Features Hopfield Networks generalize retrieval to previously unseen examples
por: Kalaj, Silvio, et al.
Publicado: (2024)
por: Kalaj, Silvio, et al.
Publicado: (2024)
Single-Head Attention in High Dimensions: A Theory of Generalization, Weights Spectra, and Scaling Laws
por: Boncoraglio, Fabrizio, et al.
Publicado: (2025)
por: Boncoraglio, Fabrizio, et al.
Publicado: (2025)
The Quantization Model of Neural Scaling
por: Michaud, Eric J., et al.
Publicado: (2023)
por: Michaud, Eric J., et al.
Publicado: (2023)
Optimal Spectral Transitions in High-Dimensional Multi-Index Models
por: Defilippis, Leonardo, et al.
Publicado: (2025)
por: Defilippis, Leonardo, et al.
Publicado: (2025)
The Effect of Optimal Self-Distillation in Noisy Gaussian Mixture Model
por: Takanami, Kaito, et al.
Publicado: (2025)
por: Takanami, Kaito, et al.
Publicado: (2025)
EB-RANSAC: Random Sample Consensus based on Energy-Based Model
por: Yasuda, Muneki, et al.
Publicado: (2026)
por: Yasuda, Muneki, et al.
Publicado: (2026)
Learning curves theory for hierarchically compositional data with power-law distributed features
por: Cagnetta, Francesco, et al.
Publicado: (2025)
por: Cagnetta, Francesco, et al.
Publicado: (2025)
Small Singular Values Matter: A Random Matrix Analysis of Transformer Models
por: Staats, Max, et al.
Publicado: (2024)
por: Staats, Max, et al.
Publicado: (2024)
Theory of Speciation Transitions in Diffusion Models with General Class Structure
por: Achilli, Beatrice, et al.
Publicado: (2026)
por: Achilli, Beatrice, et al.
Publicado: (2026)
Statistical Physics of Deep Neural Networks: Generalization Capability, Beyond the Infinite Width, and Feature Learning
por: Ariosto, Sebastiano
Publicado: (2025)
por: Ariosto, Sebastiano
Publicado: (2025)
Random features and polynomial rules
por: Aguirre-López, Fabián, et al.
Publicado: (2024)
por: Aguirre-López, Fabián, et al.
Publicado: (2024)
Optimal thresholds and algorithms for a model of multi-modal learning in high dimensions
por: Keup, Christian, et al.
Publicado: (2024)
por: Keup, Christian, et al.
Publicado: (2024)
Ejemplares similares
-
How Feature Learning Can Improve Neural Scaling Laws
por: Bordelon, Blake, et al.
Publicado: (2024) -
A Dynamical Model of Neural Scaling Laws
por: Bordelon, Blake, et al.
Publicado: (2024) -
Theory of Scaling Laws for In-Context Regression: Depth, Width, Context and Time
por: Bordelon, Blake, et al.
Publicado: (2025) -
Disordered Dynamics in High Dimensions: Connections to Random Matrices and Machine Learning
por: Bordelon, Blake, et al.
Publicado: (2026) -
Transfer Learning in Infinite Width Feature Learning Networks
por: Lauditi, Clarissa, et al.
Publicado: (2025)