A Theoretical Framework for Grokking: Interpolation followed by Riemannian Norm Minimisation
Fuente:
arXiv
Saved in:
| Main Authors: | Boursier, Etienne, Pesme, Scott, Dragomir, Radu-Alexandru |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Implicit Bias of Mirror Flow on Separable Data
by: Pesme, Scott, et al.
Published: (2024)
by: Pesme, Scott, et al.
Published: (2024)
Leveraging Continuous Time to Understand Momentum When Training Diagonal Linear Networks
by: Papazov, Hristo, et al.
Published: (2024)
by: Papazov, Hristo, et al.
Published: (2024)
MAP Estimation with Denoisers: Convergence Rates and Guarantees
by: Pesme, Scott, et al.
Published: (2025)
by: Pesme, Scott, et al.
Published: (2025)
Minimisation of Polyak-Łojasewicz Functions Using Random Zeroth-Order Oracles
by: Farzin, Amir Ali, et al.
Published: (2024)
by: Farzin, Amir Ali, et al.
Published: (2024)
A Framework for Bilevel Optimization on Riemannian Manifolds
by: Han, Andi, et al.
Published: (2024)
by: Han, Andi, et al.
Published: (2024)
On the Mechanism and Dynamics of Modular Addition: Fourier Features, Lottery Ticket, and Grokking
by: He, Jianliang, et al.
Published: (2026)
by: He, Jianliang, et al.
Published: (2026)
Consensus-Based Optimization Beyond Finite-Time Analysis
by: Bianchi, Pascal, et al.
Published: (2025)
by: Bianchi, Pascal, et al.
Published: (2025)
Importance Sampling Optimization with Laplace Principle
by: Dragomir, Radu-Alexandru, et al.
Published: (2026)
by: Dragomir, Radu-Alexandru, et al.
Published: (2026)
Convex quartic problems: homogenized gradient method and preconditioning
by: Dragomir, Radu-Alexandru, et al.
Published: (2023)
by: Dragomir, Radu-Alexandru, et al.
Published: (2023)
Preconditioned Norms: A Unified Framework for Steepest Descent, Quasi-Newton and Adaptive Methods
by: Veprikov, Andrey, et al.
Published: (2025)
by: Veprikov, Andrey, et al.
Published: (2025)
Minimisation of Submodular Functions Using Gaussian Zeroth-Order Random Oracles
by: Farzin, Amir Ali, et al.
Published: (2025)
by: Farzin, Amir Ali, et al.
Published: (2025)
Clipping Improves Adam-Norm and AdaGrad-Norm when the Noise Is Heavy-Tailed
by: Chezhegov, Savelii, et al.
Published: (2024)
by: Chezhegov, Savelii, et al.
Published: (2024)
Riemannian Dueling Optimization
by: Ren, Yuxuan, et al.
Published: (2026)
by: Ren, Yuxuan, et al.
Published: (2026)
Grokking or Glitching? How Low-Precision Drives Slingshot Loss Spikes
by: Hanqing, Liu, et al.
Published: (2026)
by: Hanqing, Liu, et al.
Published: (2026)
Old Optimizer, New Norm: An Anthology
by: Bernstein, Jeremy, et al.
Published: (2024)
by: Bernstein, Jeremy, et al.
Published: (2024)
Tame Riemannian Stochastic Approximation
by: Aspman, Johannes, et al.
Published: (2023)
by: Aspman, Johannes, et al.
Published: (2023)
Riemannian Neural Optimal Transport
by: Micheli, Alessandro, et al.
Published: (2026)
by: Micheli, Alessandro, et al.
Published: (2026)
Mirror Descent on Riemannian Manifolds
by: Jiang, Jiaxin, et al.
Published: (2026)
by: Jiang, Jiaxin, et al.
Published: (2026)
A Control Theoretic Framework for Adaptive Gradient Optimizers in Machine Learning
by: Chakrabarti, Kushal, et al.
Published: (2022)
by: Chakrabarti, Kushal, et al.
Published: (2022)
Bernoulli-LoRA: A Theoretical Framework for Randomized Low-Rank Adaptation
by: Sokolov, Igor, et al.
Published: (2025)
by: Sokolov, Igor, et al.
Published: (2025)
Non-Singularity of the Gradient Descent map for Neural Networks with Piecewise Analytic Activations
by: Crăciun, Alexandru, et al.
Published: (2025)
by: Crăciun, Alexandru, et al.
Published: (2025)
On the Interpolation Effect of Score Smoothing in Diffusion Models
by: Chen, Zhengdao
Published: (2025)
by: Chen, Zhengdao
Published: (2025)
Muon Optimizes Under Spectral Norm Constraints
by: Chen, Lizhang, et al.
Published: (2025)
by: Chen, Lizhang, et al.
Published: (2025)
Stable Nonconvex-Nonconcave Training via Linear Interpolation
by: Pethick, Thomas, et al.
Published: (2023)
by: Pethick, Thomas, et al.
Published: (2023)
Universal Architectures for the Learning of Polyhedral Norms and Convex Regularizers
by: Unser, Michael, et al.
Published: (2025)
by: Unser, Michael, et al.
Published: (2025)
Training Deep Learning Models with Norm-Constrained LMOs
by: Pethick, Thomas, et al.
Published: (2025)
by: Pethick, Thomas, et al.
Published: (2025)
Small Gradient Norm Regret for Online Convex Optimization
by: Gao, Wenzhi, et al.
Published: (2026)
by: Gao, Wenzhi, et al.
Published: (2026)
Landing with the Score: Riemannian Optimization through Denoising
by: Kharitenko, Andrey, et al.
Published: (2025)
by: Kharitenko, Andrey, et al.
Published: (2025)
Riemannian coordinate descent algorithms on matrix manifolds
by: Han, Andi, et al.
Published: (2024)
by: Han, Andi, et al.
Published: (2024)
Fast Last-Iterate Convergence of SGD in the Smooth Interpolation Regime
by: Attia, Amit, et al.
Published: (2025)
by: Attia, Amit, et al.
Published: (2025)
General Loss Functions Lead to (Approximate) Interpolation in High Dimensions
by: Lai, Kuo-Wei, et al.
Published: (2023)
by: Lai, Kuo-Wei, et al.
Published: (2023)
Faster Convergence of Stochastic Accelerated Gradient Descent under Interpolation
by: Mishkin, Aaron, et al.
Published: (2024)
by: Mishkin, Aaron, et al.
Published: (2024)
Fast, Accurate Manifold Denoising by Tunneling Riemannian Optimization
by: Wang, Shiyu, et al.
Published: (2025)
by: Wang, Shiyu, et al.
Published: (2025)
Implicit Riemannian Optimism with Applications to Min-Max Problems
by: Roux, Christophe, et al.
Published: (2025)
by: Roux, Christophe, et al.
Published: (2025)
Extragradient Type Methods for Riemannian Variational Inequality Problems
by: Hu, Zihao, et al.
Published: (2023)
by: Hu, Zihao, et al.
Published: (2023)
Riemannian Stochastic Gradient Method for Nested Composition Optimization
by: Zhang, Dewei, et al.
Published: (2022)
by: Zhang, Dewei, et al.
Published: (2022)
Riemannian Optimization for Hadamard Products of Low-Rank Matrices
by: Jawanpuria, Pratik, et al.
Published: (2026)
by: Jawanpuria, Pratik, et al.
Published: (2026)
Gradient Clipping Beyond Vector Norms: A Spectral Approach for Matrix-Valued Parameters
by: Yukhimchuk, Alexander, et al.
Published: (2026)
by: Yukhimchuk, Alexander, et al.
Published: (2026)
Interpolation Conditions for Data Consistency and Prediction in Noisy Linear Systems
by: Vanelli, Martina, et al.
Published: (2025)
by: Vanelli, Martina, et al.
Published: (2025)
Federated Learning on Riemannian Manifolds: A Gradient-Free Projection-Based Approach
by: Wang, Hongye, et al.
Published: (2025)
by: Wang, Hongye, et al.
Published: (2025)
Similar Items
-
Implicit Bias of Mirror Flow on Separable Data
by: Pesme, Scott, et al.
Published: (2024) -
Leveraging Continuous Time to Understand Momentum When Training Diagonal Linear Networks
by: Papazov, Hristo, et al.
Published: (2024) -
MAP Estimation with Denoisers: Convergence Rates and Guarantees
by: Pesme, Scott, et al.
Published: (2025) -
Minimisation of Polyak-Łojasewicz Functions Using Random Zeroth-Order Oracles
by: Farzin, Amir Ali, et al.
Published: (2024) -
A Framework for Bilevel Optimization on Riemannian Manifolds
by: Han, Andi, et al.
Published: (2024)