Deep Linear Network Training Dynamics from Random Initialization: Data, Width, Depth, and Hyperparameter Transfer
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Bordelon, Blake, Pehlevan, Cengiz |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Transfer Learning in Infinite Width Feature Learning Networks
von: Lauditi, Clarissa, et al.
Veröffentlicht: (2025)
von: Lauditi, Clarissa, et al.
Veröffentlicht: (2025)
Disordered Dynamics in High Dimensions: Connections to Random Matrices and Machine Learning
von: Bordelon, Blake, et al.
Veröffentlicht: (2026)
von: Bordelon, Blake, et al.
Veröffentlicht: (2026)
Theory of Scaling Laws for In-Context Regression: Depth, Width, Context and Time
von: Bordelon, Blake, et al.
Veröffentlicht: (2025)
von: Bordelon, Blake, et al.
Veröffentlicht: (2025)
Spectral Dynamics in Deep Networks: Feature Learning, Outlier Escape, and Learning Rate Transfer
von: Lauditi, Clarissa, et al.
Veröffentlicht: (2026)
von: Lauditi, Clarissa, et al.
Veröffentlicht: (2026)
A Dynamical Model of Neural Scaling Laws
von: Bordelon, Blake, et al.
Veröffentlicht: (2024)
von: Bordelon, Blake, et al.
Veröffentlicht: (2024)
Grokking as the Transition from Lazy to Rich Training Dynamics
von: Kumar, Tanishq, et al.
Veröffentlicht: (2023)
von: Kumar, Tanishq, et al.
Veröffentlicht: (2023)
Adaptive kernel predictors from feature-learning infinite limits of neural networks
von: Lauditi, Clarissa, et al.
Veröffentlicht: (2025)
von: Lauditi, Clarissa, et al.
Veröffentlicht: (2025)
Infinite Limits of Multi-head Transformer Dynamics
von: Bordelon, Blake, et al.
Veröffentlicht: (2024)
von: Bordelon, Blake, et al.
Veröffentlicht: (2024)
How Feature Learning Can Improve Neural Scaling Laws
von: Bordelon, Blake, et al.
Veröffentlicht: (2024)
von: Bordelon, Blake, et al.
Veröffentlicht: (2024)
Two-Point Deterministic Equivalence for Stochastic Gradient Dynamics in Linear Models
von: Atanasov, Alexander, et al.
Veröffentlicht: (2025)
von: Atanasov, Alexander, et al.
Veröffentlicht: (2025)
Dynamically Learning to Integrate in Recurrent Neural Networks
von: Bordelon, Blake, et al.
Veröffentlicht: (2025)
von: Bordelon, Blake, et al.
Veröffentlicht: (2025)
Theory of Optimal Learning Rate Schedules and Scaling Laws for a Random Feature Model
von: Bordelon, Blake, et al.
Veröffentlicht: (2026)
von: Bordelon, Blake, et al.
Veröffentlicht: (2026)
Nadaraya-Watson kernel smoothing as a random energy model
von: Zavatone-Veth, Jacob A., et al.
Veröffentlicht: (2024)
von: Zavatone-Veth, Jacob A., et al.
Veröffentlicht: (2024)
No Free Lunch From Random Feature Ensembles: Scaling Laws and Near-Optimality Conditions
von: Ruben, Benjamin S., et al.
Veröffentlicht: (2024)
von: Ruben, Benjamin S., et al.
Veröffentlicht: (2024)
A solvable model of learning generative diffusion: theory and insights
von: Cui, Hugo, et al.
Veröffentlicht: (2025)
von: Cui, Hugo, et al.
Veröffentlicht: (2025)
Learning Curves for Noisy Heterogeneous Feature-Subsampled Ridge Ensembles
von: Ruben, Benjamin S., et al.
Veröffentlicht: (2023)
von: Ruben, Benjamin S., et al.
Veröffentlicht: (2023)
Risk and cross validation in ridge regression with correlated samples
von: Atanasov, Alexander, et al.
Veröffentlicht: (2024)
von: Atanasov, Alexander, et al.
Veröffentlicht: (2024)
Scaling and renormalization in high-dimensional regression
von: Atanasov, Alexander, et al.
Veröffentlicht: (2024)
von: Atanasov, Alexander, et al.
Veröffentlicht: (2024)
How does training shape the Riemannian geometry of neural network representations?
von: Zavatone-Veth, Jacob A., et al.
Veröffentlicht: (2023)
von: Zavatone-Veth, Jacob A., et al.
Veröffentlicht: (2023)
Asymptotic theory of in-context learning by linear attention
von: Lu, Yue M., et al.
Veröffentlicht: (2024)
von: Lu, Yue M., et al.
Veröffentlicht: (2024)
A note on the dynamics of extended-context disordered kinetic spin models
von: Zavatone-Veth, Jacob A., et al.
Veröffentlicht: (2025)
von: Zavatone-Veth, Jacob A., et al.
Veröffentlicht: (2025)
Statistical Physics of Deep Neural Networks: Generalization Capability, Beyond the Infinite Width, and Feature Learning
von: Ariosto, Sebastiano
Veröffentlicht: (2025)
von: Ariosto, Sebastiano
Veröffentlicht: (2025)
A Random-Matrix Criterion for Initializing Gated Recurrent Neural Networks
von: Fioratti, Tommaso, et al.
Veröffentlicht: (2026)
von: Fioratti, Tommaso, et al.
Veröffentlicht: (2026)
How Deep Networks Learn Sparse and Hierarchical Data: the Sparse Random Hierarchy Model
von: Tomasini, Umberto, et al.
Veröffentlicht: (2024)
von: Tomasini, Umberto, et al.
Veröffentlicht: (2024)
Quantifying Hyperparameter Transfer and the Importance of Embedding Layer Learning Rate
von: Kalra, Dayal Singh, et al.
Veröffentlicht: (2026)
von: Kalra, Dayal Singh, et al.
Veröffentlicht: (2026)
Precise Dynamics of Diagonal Linear Networks: A Unifying Analysis by Dynamical Mean-Field Theory
von: Nishiyama, Sota, et al.
Veröffentlicht: (2025)
von: Nishiyama, Sota, et al.
Veröffentlicht: (2025)
The Training Process of Many Deep Networks Explores the Same Low-Dimensional Manifold
von: Mao, Jialin, et al.
Veröffentlicht: (2023)
von: Mao, Jialin, et al.
Veröffentlicht: (2023)
Controlled Langevin Dynamics for Sampling of Feedforward Neural Networks Trained with Minibatches
von: Zambon, Alessandro, et al.
Veröffentlicht: (2026)
von: Zambon, Alessandro, et al.
Veröffentlicht: (2026)
Initial Guessing Bias: How Untrained Networks Favor Some Classes
von: Francazi, Emanuele, et al.
Veröffentlicht: (2023)
von: Francazi, Emanuele, et al.
Veröffentlicht: (2023)
Exact Learning Dynamics of In-Context Learning in Linear Transformers and Its Application to Non-Linear Transformers
von: Mainali, Nischal, et al.
Veröffentlicht: (2025)
von: Mainali, Nischal, et al.
Veröffentlicht: (2025)
Transfer Learning in $\ell_1$ Regularized Regression: Hyperparameter Selection Strategy based on Sharp Asymptotic Analysis
von: Okajima, Koki, et al.
Veröffentlicht: (2024)
von: Okajima, Koki, et al.
Veröffentlicht: (2024)
Bias in Motion: Theoretical Insights into the Dynamics of Bias in SGD Training
von: Jain, Anchit, et al.
Veröffentlicht: (2024)
von: Jain, Anchit, et al.
Veröffentlicht: (2024)
Dataset-Free Weight-Initialization on Restricted Boltzmann Machine
von: Yasuda, Muneki, et al.
Veröffentlicht: (2024)
von: Yasuda, Muneki, et al.
Veröffentlicht: (2024)
High-Dimensional Analysis of Gradient Flow for Extensive-Width Quadratic Neural Networks
von: Martin, Simon, et al.
Veröffentlicht: (2026)
von: Martin, Simon, et al.
Veröffentlicht: (2026)
Training Dynamics of Nonlinear Contrastive Learning Model in the High Dimensional Limit
von: Meng, Lineghuan, et al.
Veröffentlicht: (2024)
von: Meng, Lineghuan, et al.
Veröffentlicht: (2024)
A Theory of Saddle Escape in Deep Nonlinear Networks
von: Rawal, Divit, et al.
Veröffentlicht: (2026)
von: Rawal, Divit, et al.
Veröffentlicht: (2026)
Why Diffusion Models Don't Memorize: The Role of Implicit Dynamical Regularization in Training
von: Bonnaire, Tony, et al.
Veröffentlicht: (2025)
von: Bonnaire, Tony, et al.
Veröffentlicht: (2025)
Where You Place the Norm Matters: From Prejudiced to Neutral Initializations
von: Francazi, Emanuele, et al.
Veröffentlicht: (2025)
von: Francazi, Emanuele, et al.
Veröffentlicht: (2025)
The Interplay of Data Structure and Imbalance in the Learning Dynamics of Diffusion Models
von: Nicoletti, Flavio, et al.
Veröffentlicht: (2026)
von: Nicoletti, Flavio, et al.
Veröffentlicht: (2026)
Growing Neural Networks: Dynamic Evolution through Gradient Descent
von: Radhakrishnan, Anil, et al.
Veröffentlicht: (2025)
von: Radhakrishnan, Anil, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Transfer Learning in Infinite Width Feature Learning Networks
von: Lauditi, Clarissa, et al.
Veröffentlicht: (2025) -
Disordered Dynamics in High Dimensions: Connections to Random Matrices and Machine Learning
von: Bordelon, Blake, et al.
Veröffentlicht: (2026) -
Theory of Scaling Laws for In-Context Regression: Depth, Width, Context and Time
von: Bordelon, Blake, et al.
Veröffentlicht: (2025) -
Spectral Dynamics in Deep Networks: Feature Learning, Outlier Escape, and Learning Rate Transfer
von: Lauditi, Clarissa, et al.
Veröffentlicht: (2026) -
A Dynamical Model of Neural Scaling Laws
von: Bordelon, Blake, et al.
Veröffentlicht: (2024)