Grokking as a First Order Phase Transition in Two Layer Networks
Fuente:
arXiv
Saved in:
| Main Authors: | Rubin, Noa, Seroussi, Inbar, Ringel, Zohar |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Applications of Statistical Field Theory in Deep Learning
by: Ringel, Zohar, et al.
Published: (2025)
by: Ringel, Zohar, et al.
Published: (2025)
From Kernels to Features: A Multi-Scale Adaptive Theory of Feature Learning
by: Rubin, Noa, et al.
Published: (2025)
by: Rubin, Noa, et al.
Published: (2025)
Algorithmic Task Capture, Computational Complexity, and Inductive Bias of Infinite Transformers
by: Davidovich, Orit, et al.
Published: (2026)
by: Davidovich, Orit, et al.
Published: (2026)
Demystifying Spectral Bias on Real-World Data
by: Lavie, Itay, et al.
Published: (2024)
by: Lavie, Itay, et al.
Published: (2024)
Towards Understanding Inductive Bias in Transformers: A View From Infinity
by: Lavie, Itay, et al.
Published: (2024)
by: Lavie, Itay, et al.
Published: (2024)
Grokking as the Transition from Lazy to Rich Training Dynamics
by: Kumar, Tanishq, et al.
Published: (2023)
by: Kumar, Tanishq, et al.
Published: (2023)
Grokking as Dimensional Phase Transition in Neural Networks
by: Wang, Ping
Published: (2026)
by: Wang, Ping
Published: (2026)
Is Grokking a Computational Glass Relaxation?
by: Zhang, Xiaotian, et al.
Published: (2025)
by: Zhang, Xiaotian, et al.
Published: (2025)
Wilsonian Renormalization of Neural Network Gaussian Processes
by: Howard, Jessica N., et al.
Published: (2024)
by: Howard, Jessica N., et al.
Published: (2024)
Renormalization group for deep neural networks: Universality of learning and scaling laws
by: Coppola, Gorka Peraza, et al.
Published: (2025)
by: Coppola, Gorka Peraza, et al.
Published: (2025)
Dynamical Decoupling of Generalization and Overfitting in Large Two-Layer Networks
by: Montanari, Andrea, et al.
Published: (2025)
by: Montanari, Andrea, et al.
Published: (2025)
Lecture notes: From Gaussian processes to feature learning
by: Helias, Moritz, et al.
Published: (2026)
by: Helias, Moritz, et al.
Published: (2026)
Grokking at the Edge of Linear Separability
by: Beck, Alon, et al.
Published: (2024)
by: Beck, Alon, et al.
Published: (2024)
Critical Phase Transition in Large Language Models
by: Nakaishi, Kai, et al.
Published: (2024)
by: Nakaishi, Kai, et al.
Published: (2024)
Grokking vs. Learning: Same Features, Different Encodings
by: Manning-Coe, Dmitry, et al.
Published: (2025)
by: Manning-Coe, Dmitry, et al.
Published: (2025)
Geometric Entropy and Retrieval Phase Transitions in Continuous Thermal Dense Associative Memory
by: Petrova, Tatiana, et al.
Published: (2026)
by: Petrova, Tatiana, et al.
Published: (2026)
Critical feature learning in deep neural networks
by: Fischer, Kirsten, et al.
Published: (2024)
by: Fischer, Kirsten, et al.
Published: (2024)
Grokking in Linear Estimators -- A Solvable Model that Groks without Understanding
by: Levi, Noam, et al.
Published: (2023)
by: Levi, Noam, et al.
Published: (2023)
Asymptotics of SGD in Sequence-Single Index Models and Single-Layer Attention Networks
by: Arnaboldi, Luca, et al.
Published: (2025)
by: Arnaboldi, Luca, et al.
Published: (2025)
Rigorous Asymptotics for First-Order Algorithms Through the Dynamical Cavity Method
by: Dandi, Yatin, et al.
Published: (2026)
by: Dandi, Yatin, et al.
Published: (2026)
Grokking in the Ising Model
by: Hutchison, Karolina, et al.
Published: (2025)
by: Hutchison, Karolina, et al.
Published: (2025)
Universal Scaling Laws of Absorbing Phase Transitions in Artificial Deep Neural Networks
by: Tamai, Keiichi, et al.
Published: (2023)
by: Tamai, Keiichi, et al.
Published: (2023)
Attention to Order: Transformers Discover Phase Transitions via Learnability
by: Özönder, Şener
Published: (2025)
by: Özönder, Şener
Published: (2025)
Hebbian Learning from First Principles
by: Albanese, Linda, et al.
Published: (2024)
by: Albanese, Linda, et al.
Published: (2024)
Optimal Spectral Transitions in High-Dimensional Multi-Index Models
by: Defilippis, Leonardo, et al.
Published: (2025)
by: Defilippis, Leonardo, et al.
Published: (2025)
Theory of Speciation Transitions in Diffusion Models with General Class Structure
by: Achilli, Beatrice, et al.
Published: (2026)
by: Achilli, Beatrice, et al.
Published: (2026)
A Boundary-Layer Mechanism for One-Third Scaling in Online Softmax Classification
by: Kühn, Marcel, et al.
Published: (2026)
by: Kühn, Marcel, et al.
Published: (2026)
Statistical physics through the lens of real-space mutual information
by: Gökmen, Doruk Efe, et al.
Published: (2021)
by: Gökmen, Doruk Efe, et al.
Published: (2021)
Dimensional Criticality at Grokking Across MLPs and Transformers
by: Wang, Ping
Published: (2026)
by: Wang, Ping
Published: (2026)
Two-Point Deterministic Equivalence for Stochastic Gradient Dynamics in Linear Models
by: Atanasov, Alexander, et al.
Published: (2025)
by: Atanasov, Alexander, et al.
Published: (2025)
Inferring Higher-Order Couplings with Neural Networks
by: Decelle, Aurélien, et al.
Published: (2025)
by: Decelle, Aurélien, et al.
Published: (2025)
Grokking Modular Polynomials
by: Doshi, Darshil, et al.
Published: (2024)
by: Doshi, Darshil, et al.
Published: (2024)
Formation of Representations in Neural Networks
by: Ziyin, Liu, et al.
Published: (2024)
by: Ziyin, Liu, et al.
Published: (2024)
Graph Neural Networks Do Not Always Oversmooth
by: Epping, Bastian, et al.
Published: (2024)
by: Epping, Bastian, et al.
Published: (2024)
Transfer Learning in Infinite Width Feature Learning Networks
by: Lauditi, Clarissa, et al.
Published: (2025)
by: Lauditi, Clarissa, et al.
Published: (2025)
A Theory of Saddle Escape in Deep Nonlinear Networks
by: Rawal, Divit, et al.
Published: (2026)
by: Rawal, Divit, et al.
Published: (2026)
From Classical to Quantum: Extending Prometheus for Unsupervised Discovery of Phase Transitions in Three Dimensions and Quantum Systems
by: Yee, Brandon, et al.
Published: (2026)
by: Yee, Brandon, et al.
Published: (2026)
Demolition and Reinforcement of Memories in Spin-Glass-like Neural Networks
by: Ventura, Enrico
Published: (2024)
by: Ventura, Enrico
Published: (2024)
Growing Neural Networks: Dynamic Evolution through Gradient Descent
by: Radhakrishnan, Anil, et al.
Published: (2025)
by: Radhakrishnan, Anil, et al.
Published: (2025)
The Rules-and-Facts Model for Simultaneous Generalization and Memorization in Neural Networks
by: Farné, Gabriele, et al.
Published: (2026)
by: Farné, Gabriele, et al.
Published: (2026)
Similar Items
-
Applications of Statistical Field Theory in Deep Learning
by: Ringel, Zohar, et al.
Published: (2025) -
From Kernels to Features: A Multi-Scale Adaptive Theory of Feature Learning
by: Rubin, Noa, et al.
Published: (2025) -
Algorithmic Task Capture, Computational Complexity, and Inductive Bias of Infinite Transformers
by: Davidovich, Orit, et al.
Published: (2026) -
Demystifying Spectral Bias on Real-World Data
by: Lavie, Itay, et al.
Published: (2024) -
Towards Understanding Inductive Bias in Transformers: A View From Infinity
by: Lavie, Itay, et al.
Published: (2024)