Grokking as the Transition from Lazy to Rich Training Dynamics
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kumar, Tanishq, Bordelon, Blake, Gershman, Samuel J., Pehlevan, Cengiz |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Deep Linear Network Training Dynamics from Random Initialization: Data, Width, Depth, and Hyperparameter Transfer
von: Bordelon, Blake, et al.
Veröffentlicht: (2025)
von: Bordelon, Blake, et al.
Veröffentlicht: (2025)
Disordered Dynamics in High Dimensions: Connections to Random Matrices and Machine Learning
von: Bordelon, Blake, et al.
Veröffentlicht: (2026)
von: Bordelon, Blake, et al.
Veröffentlicht: (2026)
A Dynamical Model of Neural Scaling Laws
von: Bordelon, Blake, et al.
Veröffentlicht: (2024)
von: Bordelon, Blake, et al.
Veröffentlicht: (2024)
Adaptive kernel predictors from feature-learning infinite limits of neural networks
von: Lauditi, Clarissa, et al.
Veröffentlicht: (2025)
von: Lauditi, Clarissa, et al.
Veröffentlicht: (2025)
Infinite Limits of Multi-head Transformer Dynamics
von: Bordelon, Blake, et al.
Veröffentlicht: (2024)
von: Bordelon, Blake, et al.
Veröffentlicht: (2024)
How Feature Learning Can Improve Neural Scaling Laws
von: Bordelon, Blake, et al.
Veröffentlicht: (2024)
von: Bordelon, Blake, et al.
Veröffentlicht: (2024)
Transfer Learning in Infinite Width Feature Learning Networks
von: Lauditi, Clarissa, et al.
Veröffentlicht: (2025)
von: Lauditi, Clarissa, et al.
Veröffentlicht: (2025)
Spectral Dynamics in Deep Networks: Feature Learning, Outlier Escape, and Learning Rate Transfer
von: Lauditi, Clarissa, et al.
Veröffentlicht: (2026)
von: Lauditi, Clarissa, et al.
Veröffentlicht: (2026)
Theory of Scaling Laws for In-Context Regression: Depth, Width, Context and Time
von: Bordelon, Blake, et al.
Veröffentlicht: (2025)
von: Bordelon, Blake, et al.
Veröffentlicht: (2025)
Two-Point Deterministic Equivalence for Stochastic Gradient Dynamics in Linear Models
von: Atanasov, Alexander, et al.
Veröffentlicht: (2025)
von: Atanasov, Alexander, et al.
Veröffentlicht: (2025)
Dynamically Learning to Integrate in Recurrent Neural Networks
von: Bordelon, Blake, et al.
Veröffentlicht: (2025)
von: Bordelon, Blake, et al.
Veröffentlicht: (2025)
Theory of Optimal Learning Rate Schedules and Scaling Laws for a Random Feature Model
von: Bordelon, Blake, et al.
Veröffentlicht: (2026)
von: Bordelon, Blake, et al.
Veröffentlicht: (2026)
Nadaraya-Watson kernel smoothing as a random energy model
von: Zavatone-Veth, Jacob A., et al.
Veröffentlicht: (2024)
von: Zavatone-Veth, Jacob A., et al.
Veröffentlicht: (2024)
A solvable model of learning generative diffusion: theory and insights
von: Cui, Hugo, et al.
Veröffentlicht: (2025)
von: Cui, Hugo, et al.
Veröffentlicht: (2025)
Learning Curves for Noisy Heterogeneous Feature-Subsampled Ridge Ensembles
von: Ruben, Benjamin S., et al.
Veröffentlicht: (2023)
von: Ruben, Benjamin S., et al.
Veröffentlicht: (2023)
Risk and cross validation in ridge regression with correlated samples
von: Atanasov, Alexander, et al.
Veröffentlicht: (2024)
von: Atanasov, Alexander, et al.
Veröffentlicht: (2024)
Scaling and renormalization in high-dimensional regression
von: Atanasov, Alexander, et al.
Veröffentlicht: (2024)
von: Atanasov, Alexander, et al.
Veröffentlicht: (2024)
No Free Lunch From Random Feature Ensembles: Scaling Laws and Near-Optimality Conditions
von: Ruben, Benjamin S., et al.
Veröffentlicht: (2024)
von: Ruben, Benjamin S., et al.
Veröffentlicht: (2024)
How does training shape the Riemannian geometry of neural network representations?
von: Zavatone-Veth, Jacob A., et al.
Veröffentlicht: (2023)
von: Zavatone-Veth, Jacob A., et al.
Veröffentlicht: (2023)
Grokking as a First Order Phase Transition in Two Layer Networks
von: Rubin, Noa, et al.
Veröffentlicht: (2023)
von: Rubin, Noa, et al.
Veröffentlicht: (2023)
Asymptotic theory of in-context learning by linear attention
von: Lu, Yue M., et al.
Veröffentlicht: (2024)
von: Lu, Yue M., et al.
Veröffentlicht: (2024)
Is Grokking a Computational Glass Relaxation?
von: Zhang, Xiaotian, et al.
Veröffentlicht: (2025)
von: Zhang, Xiaotian, et al.
Veröffentlicht: (2025)
A note on the dynamics of extended-context disordered kinetic spin models
von: Zavatone-Veth, Jacob A., et al.
Veröffentlicht: (2025)
von: Zavatone-Veth, Jacob A., et al.
Veröffentlicht: (2025)
Grokking at the Edge of Linear Separability
von: Beck, Alon, et al.
Veröffentlicht: (2024)
von: Beck, Alon, et al.
Veröffentlicht: (2024)
Grokking as Dimensional Phase Transition in Neural Networks
von: Wang, Ping
Veröffentlicht: (2026)
von: Wang, Ping
Veröffentlicht: (2026)
Grokking vs. Learning: Same Features, Different Encodings
von: Manning-Coe, Dmitry, et al.
Veröffentlicht: (2025)
von: Manning-Coe, Dmitry, et al.
Veröffentlicht: (2025)
Grokking in Linear Estimators -- A Solvable Model that Groks without Understanding
von: Levi, Noam, et al.
Veröffentlicht: (2023)
von: Levi, Noam, et al.
Veröffentlicht: (2023)
Bias in Motion: Theoretical Insights into the Dynamics of Bias in SGD Training
von: Jain, Anchit, et al.
Veröffentlicht: (2024)
von: Jain, Anchit, et al.
Veröffentlicht: (2024)
Training Dynamics of Nonlinear Contrastive Learning Model in the High Dimensional Limit
von: Meng, Lineghuan, et al.
Veröffentlicht: (2024)
von: Meng, Lineghuan, et al.
Veröffentlicht: (2024)
Controlled Langevin Dynamics for Sampling of Feedforward Neural Networks Trained with Minibatches
von: Zambon, Alessandro, et al.
Veröffentlicht: (2026)
von: Zambon, Alessandro, et al.
Veröffentlicht: (2026)
Grokking in the Ising Model
von: Hutchison, Karolina, et al.
Veröffentlicht: (2025)
von: Hutchison, Karolina, et al.
Veröffentlicht: (2025)
Why Diffusion Models Don't Memorize: The Role of Implicit Dynamical Regularization in Training
von: Bonnaire, Tony, et al.
Veröffentlicht: (2025)
von: Bonnaire, Tony, et al.
Veröffentlicht: (2025)
Critical Phase Transition in Large Language Models
von: Nakaishi, Kai, et al.
Veröffentlicht: (2024)
von: Nakaishi, Kai, et al.
Veröffentlicht: (2024)
Optimal Spectral Transitions in High-Dimensional Multi-Index Models
von: Defilippis, Leonardo, et al.
Veröffentlicht: (2025)
von: Defilippis, Leonardo, et al.
Veröffentlicht: (2025)
Theory of Speciation Transitions in Diffusion Models with General Class Structure
von: Achilli, Beatrice, et al.
Veröffentlicht: (2026)
von: Achilli, Beatrice, et al.
Veröffentlicht: (2026)
Geometric Entropy and Retrieval Phase Transitions in Continuous Thermal Dense Associative Memory
von: Petrova, Tatiana, et al.
Veröffentlicht: (2026)
von: Petrova, Tatiana, et al.
Veröffentlicht: (2026)
Training neural networks with structured noise improves classification and generalization
von: Benedetti, Marco, et al.
Veröffentlicht: (2023)
von: Benedetti, Marco, et al.
Veröffentlicht: (2023)
Dimensional Criticality at Grokking Across MLPs and Transformers
von: Wang, Ping
Veröffentlicht: (2026)
von: Wang, Ping
Veröffentlicht: (2026)
The Training Process of Many Deep Networks Explores the Same Low-Dimensional Manifold
von: Mao, Jialin, et al.
Veröffentlicht: (2023)
von: Mao, Jialin, et al.
Veröffentlicht: (2023)
A High Dimensional Statistical Model for Adversarial Training: Geometry and Trade-Offs
von: Tanner, Kasimir, et al.
Veröffentlicht: (2024)
von: Tanner, Kasimir, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Deep Linear Network Training Dynamics from Random Initialization: Data, Width, Depth, and Hyperparameter Transfer
von: Bordelon, Blake, et al.
Veröffentlicht: (2025) -
Disordered Dynamics in High Dimensions: Connections to Random Matrices and Machine Learning
von: Bordelon, Blake, et al.
Veröffentlicht: (2026) -
A Dynamical Model of Neural Scaling Laws
von: Bordelon, Blake, et al.
Veröffentlicht: (2024) -
Adaptive kernel predictors from feature-learning infinite limits of neural networks
von: Lauditi, Clarissa, et al.
Veröffentlicht: (2025) -
Infinite Limits of Multi-head Transformer Dynamics
von: Bordelon, Blake, et al.
Veröffentlicht: (2024)