Emergence in non-neural models: grokking modular arithmetic via average gradient outer product
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Mallinar, Neil, Beaglehole, Daniel, Zhu, Libin, Radhakrishnan, Adityanarayanan, Pandit, Parthe, Belkin, Mikhail |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Quadratic models for understanding catapult dynamics of neural networks
par: Zhu, Libin, et autres
Publié: (2022)
par: Zhu, Libin, et autres
Publié: (2022)
xRFM: Accurate, scalable, and interpretable feature learning models for tabular data
par: Beaglehole, Daniel, et autres
Publié: (2025)
par: Beaglehole, Daniel, et autres
Publié: (2025)
Average gradient outer product as a mechanism for deep neural collapse
par: Beaglehole, Daniel, et autres
Publié: (2024)
par: Beaglehole, Daniel, et autres
Publié: (2024)
Toward universal steering and monitoring of AI models
par: Beaglehole, Daniel, et autres
Publié: (2025)
par: Beaglehole, Daniel, et autres
Publié: (2025)
Eigenvectors of the De Bruijn Graph Laplacian: A Natural Basis for the Cut and Cycle Space
par: Philippakis, Anthony, et autres
Publié: (2024)
par: Philippakis, Anthony, et autres
Publié: (2024)
Catapults in SGD: spikes in the training loss and their impact on generalization through feature learning
par: Zhu, Libin, et autres
Publié: (2023)
par: Zhu, Libin, et autres
Publié: (2023)
Benign, Tempered, or Catastrophic: A Taxonomy of Overfitting
par: Mallinar, Neil, et autres
Publié: (2022)
par: Mallinar, Neil, et autres
Publié: (2022)
Contextual Linear Activation Steering of Language Models
par: Hsu, Brandon, et autres
Publié: (2026)
par: Hsu, Brandon, et autres
Publié: (2026)
Linear Recursive Feature Machines provably recover low-rank matrices
par: Radhakrishnan, Adityanarayanan, et autres
Publié: (2024)
par: Radhakrishnan, Adityanarayanan, et autres
Publié: (2024)
Mirror Descent on Reproducing Kernel Banach Spaces
par: Kumar, Akash, et autres
Publié: (2024)
par: Kumar, Akash, et autres
Publié: (2024)
Fast training of large kernel models with delayed projections
par: Abedsoltan, Amirhesam, et autres
Publié: (2024)
par: Abedsoltan, Amirhesam, et autres
Publié: (2024)
The Weight Gram Matrix Captures Sequential Feature Linearization in Deep Networks
par: Cha, Taehun, et autres
Publié: (2026)
par: Cha, Taehun, et autres
Publié: (2026)
On the Nystrom Approximation for Preconditioning in Kernel Machines
par: Abedsoltan, Amirhesam, et autres
Publié: (2023)
par: Abedsoltan, Amirhesam, et autres
Publié: (2023)
Context-Scaling versus Task-Scaling in In-Context Learning
par: Abedsoltan, Amirhesam, et autres
Publié: (2024)
par: Abedsoltan, Amirhesam, et autres
Publié: (2024)
Asymptotic convexity of wide and shallow neural networks
par: Borkar, Vivek, et autres
Publié: (2025)
par: Borkar, Vivek, et autres
Publié: (2025)
Breaking Data Symmetry is Needed For Generalization in Feature Learning Kernels
par: Bernal, Marcel Tomàs, et autres
Publié: (2026)
par: Bernal, Marcel Tomàs, et autres
Publié: (2026)
The Features at Convergence Theorem: a first-principles alternative to the Neural Feature Ansatz for how networks learn representations
par: Boix-Adsera, Enric, et autres
Publié: (2025)
par: Boix-Adsera, Enric, et autres
Publié: (2025)
Feature learning as alignment: a structural property of gradient descent in non-linear neural networks
par: Beaglehole, Daniel, et autres
Publié: (2024)
par: Beaglehole, Daniel, et autres
Publié: (2024)
Feature maps for the Laplacian kernel and its generalizations
par: Ahir, Sudhendu, et autres
Publié: (2025)
par: Ahir, Sudhendu, et autres
Publié: (2025)
Universality of Kernel Random Matrices and Kernel Regression in the Quadratic Regime
par: Pandit, Parthe, et autres
Publié: (2024)
par: Pandit, Parthe, et autres
Publié: (2024)
Efficient and accurate steering of Large Language Models through attention-guided feature learning
par: Davarmanesh, Parmida, et autres
Publié: (2026)
par: Davarmanesh, Parmida, et autres
Publié: (2026)
How to explain grokking
par: Kozyrev, S. V.
Publié: (2024)
par: Kozyrev, S. V.
Publié: (2024)
A rationale from frequency perspective for grokking in training neural network
par: Zhou, Zhangchen, et autres
Publié: (2024)
par: Zhou, Zhangchen, et autres
Publié: (2024)
Minimum-Norm Interpolation Under Covariate Shift
par: Mallinar, Neil, et autres
Publié: (2024)
par: Mallinar, Neil, et autres
Publié: (2024)
Using physics-inspired Singular Learning Theory to understand grokking & other phase transitions in modern neural networks
par: Lakkapragada, Anish
Publié: (2025)
par: Lakkapragada, Anish
Publié: (2025)
Learning to grok: Emergence of in-context learning and skill composition in modular arithmetic tasks
par: He, Tianyu, et autres
Publié: (2024)
par: He, Tianyu, et autres
Publié: (2024)
Catching rationalization in the act: detecting motivated reasoning before and after CoT via activation probing
par: Mirtaheri, Parsa, et autres
Publié: (2026)
par: Mirtaheri, Parsa, et autres
Publié: (2026)
Learning words in groups: fusion algebras, tensor ranks and grokking
par: Shutman, Maor, et autres
Publié: (2025)
par: Shutman, Maor, et autres
Publié: (2025)
Late-Stage Generalization Collapse in Grokking: Detecting anti-grokking with Weightwatcher
par: Prakash, Hari K, et autres
Publié: (2026)
par: Prakash, Hari K, et autres
Publié: (2026)
Detecting overfitting in Neural Networks during long-horizon grokking using Random Matrix Theory
par: Prakash, Hari K., et autres
Publié: (2026)
par: Prakash, Hari K., et autres
Publié: (2026)
General and Efficient Steering of Unconditional Diffusion
par: Wang, Qingsong, et autres
Publié: (2026)
par: Wang, Qingsong, et autres
Publié: (2026)
A Gap Between the Gaussian RKHS and Neural Networks: An Infinite-Center Asymptotic Analysis
par: Kumar, Akash, et autres
Publié: (2025)
par: Kumar, Akash, et autres
Publié: (2025)
More is Better in Modern Machine Learning: when Infinite Overparameterization is Optimal and Overfitting is Obligatory
par: Simon, James B., et autres
Publié: (2023)
par: Simon, James B., et autres
Publié: (2023)
Structure tensor Reynolds-averaged Navier-Stokes turbulence models with equivariant neural networks
par: Miller, Aaron, et autres
Publié: (2025)
par: Miller, Aaron, et autres
Publié: (2025)
Phase codes emerge in recurrent neural networks optimized for modular arithmetic
par: Murray, Keith T.
Publié: (2023)
par: Murray, Keith T.
Publié: (2023)
Unbiased least squares regression via averaged stochastic gradient descent
par: Kahalé, Nabil
Publié: (2024)
par: Kahalé, Nabil
Publié: (2024)
Average Gradient Outer Product in kernel regression provably recovers the central subspace for multi-index models
par: Zhu, Libin, et autres
Publié: (2026)
par: Zhu, Libin, et autres
Publié: (2026)
Can neural networks do arithmetic? A survey on the elementary numerical skills of state-of-the-art deep learning models
par: Testolin, Alberto
Publié: (2023)
par: Testolin, Alberto
Publié: (2023)
Multi-models with averaging in feature domain for non-invasive blood glucose estimation
par: Wei, Yiting, et autres
Publié: (2025)
par: Wei, Yiting, et autres
Publié: (2025)
Training event-based neural networks with exact gradients via Differentiable ODE Solving in JAX
par: König, Lukas, et autres
Publié: (2026)
par: König, Lukas, et autres
Publié: (2026)
Documents similaires
-
Quadratic models for understanding catapult dynamics of neural networks
par: Zhu, Libin, et autres
Publié: (2022) -
xRFM: Accurate, scalable, and interpretable feature learning models for tabular data
par: Beaglehole, Daniel, et autres
Publié: (2025) -
Average gradient outer product as a mechanism for deep neural collapse
par: Beaglehole, Daniel, et autres
Publié: (2024) -
Toward universal steering and monitoring of AI models
par: Beaglehole, Daniel, et autres
Publié: (2025) -
Eigenvectors of the De Bruijn Graph Laplacian: A Natural Basis for the Cut and Cycle Space
par: Philippakis, Anthony, et autres
Publié: (2024)