Egalitarian Gradient Descent: A Simple Approach to Accelerated Grokking
Fuente:
arXiv
Saved in:
| Main Authors: | Pasand, Ali Saheb, Dohmatob, Elvis |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
WERank: Towards Rank Degradation Prevention for Self-Supervised Learning Using Weight Regularization
by: Pasand, Ali Saheb, et al.
Published: (2024)
by: Pasand, Ali Saheb, et al.
Published: (2024)
auto-fpt: Automating Free Probability Theory Calculations for Machine Learning Theory
by: Subramonian, Arjun, et al.
Published: (2025)
by: Subramonian, Arjun, et al.
Published: (2025)
Scalable Graph Self-Supervised Learning
by: Pasand, Ali Saheb, et al.
Published: (2024)
by: Pasand, Ali Saheb, et al.
Published: (2024)
Why Less is More (Sometimes): A Theory of Data Curation
by: Dohmatob, Elvis, et al.
Published: (2025)
by: Dohmatob, Elvis, et al.
Published: (2025)
Model Collapse Demystified: The Case of Regression
by: Dohmatob, Elvis, et al.
Published: (2024)
by: Dohmatob, Elvis, et al.
Published: (2024)
Efficient Refusal Ablation in LLM through Optimal Transport
by: Nanfack, Geraldin, et al.
Published: (2026)
by: Nanfack, Geraldin, et al.
Published: (2026)
Grokfast: Accelerated Grokking by Amplifying Slow Gradients
by: Lee, Jaerin, et al.
Published: (2024)
by: Lee, Jaerin, et al.
Published: (2024)
Stable Deep Reinforcement Learning via Isotropic Gaussian Representations
by: Pasand, Ali Saheb, et al.
Published: (2026)
by: Pasand, Ali Saheb, et al.
Published: (2026)
Strong Model Collapse
by: Dohmatob, Elvis, et al.
Published: (2024)
by: Dohmatob, Elvis, et al.
Published: (2024)
NeuralGrok: Accelerate Grokking by Neural Gradient Transformation
by: Zhou, Xinyu, et al.
Published: (2025)
by: Zhou, Xinyu, et al.
Published: (2025)
REAM: Merging Improves Pruning of Experts in LLMs
by: Jha, Saurav, et al.
Published: (2026)
by: Jha, Saurav, et al.
Published: (2026)
Stacking as Accelerated Gradient Descent
by: Agarwal, Naman, et al.
Published: (2024)
by: Agarwal, Naman, et al.
Published: (2024)
An Effective Theory of Bias Amplification
by: Subramonian, Arjun, et al.
Published: (2024)
by: Subramonian, Arjun, et al.
Published: (2024)
Scaling Laws for Associative Memories
by: Cabannes, Vivien, et al.
Published: (2023)
by: Cabannes, Vivien, et al.
Published: (2023)
Beyond Model Collapse: Scaling Up with Synthesized Data Requires Verification
by: Feng, Yunzhen, et al.
Published: (2024)
by: Feng, Yunzhen, et al.
Published: (2024)
A Tale of Tails: Model Collapse as a Change of Scaling Laws
by: Dohmatob, Elvis, et al.
Published: (2024)
by: Dohmatob, Elvis, et al.
Published: (2024)
GrokAlign: Geometric Characterisation and Acceleration of Grokking
by: Walker, Thomas, et al.
Published: (2025)
by: Walker, Thomas, et al.
Published: (2025)
Anytime Acceleration of Gradient Descent
by: Zhang, Zihan, et al.
Published: (2024)
by: Zhang, Zihan, et al.
Published: (2024)
Muon Optimizer Accelerates Grokking
by: Tveit, Amund, et al.
Published: (2025)
by: Tveit, Amund, et al.
Published: (2025)
The Pitfalls of Memorization: When Memorization Hurts Generalization
by: Bayat, Reza, et al.
Published: (2024)
by: Bayat, Reza, et al.
Published: (2024)
TExplain: Explaining Learned Visual Features via Pre-trained (Frozen) Language Models
by: Taghanaki, Saeid Asgari, et al.
Published: (2023)
by: Taghanaki, Saeid Asgari, et al.
Published: (2023)
Accelerated Gradient Descent for Faster Convergence with Minimal Overhead
by: Graca, Manuel, et al.
Published: (2026)
by: Graca, Manuel, et al.
Published: (2026)
Two Speeds of Learning: A Representation-Readout Decomposition of Grokking and Double Descent
by: Chou, Chi-Ning, et al.
Published: (2026)
by: Chou, Chi-Ning, et al.
Published: (2026)
Unified View of Grokking, Double Descent and Emergent Abilities: A Perspective from Circuits Competition
by: Huang, Yufei, et al.
Published: (2024)
by: Huang, Yufei, et al.
Published: (2024)
Streaming Krylov-Accelerated Stochastic Gradient Descent
by: Thomas, Stephen
Published: (2025)
by: Thomas, Stephen
Published: (2025)
Preconditioning for Accelerated Gradient Descent Optimization and Regularization
by: Ye, Qiang
Published: (2024)
by: Ye, Qiang
Published: (2024)
Deep Learning Through A Telescoping Lens: A Simple Model Provides Empirical Insights On Grokking, Gradient Boosting & Beyond
by: Jeffares, Alan, et al.
Published: (2024)
by: Jeffares, Alan, et al.
Published: (2024)
To Grok Grokking: Provable Grokking in Ridge Regression
by: Xu, Mingyue, et al.
Published: (2026)
by: Xu, Mingyue, et al.
Published: (2026)
Large Stepsizes Accelerate Gradient Descent for Regularized Logistic Regression
by: Wu, Jingfeng, et al.
Published: (2025)
by: Wu, Jingfeng, et al.
Published: (2025)
On the Generalization of Stochastic Gradient Descent with Momentum
by: Ramezani-Kebrya, Ali, et al.
Published: (2018)
by: Ramezani-Kebrya, Ali, et al.
Published: (2018)
Accelerating Convergence of Stein Variational Gradient Descent via Deep Unfolding
by: Kawamura, Yuya, et al.
Published: (2024)
by: Kawamura, Yuya, et al.
Published: (2024)
Stochastic Bandits for Egalitarian Assignment
by: Lim, Eugene, et al.
Published: (2024)
by: Lim, Eugene, et al.
Published: (2024)
Robust Gradient Descent via Heavy-Ball Momentum with Predictive Extrapolation
by: Ali, Sarwan
Published: (2025)
by: Ali, Sarwan
Published: (2025)
Faster Convergence of Stochastic Accelerated Gradient Descent under Interpolation
by: Mishkin, Aaron, et al.
Published: (2024)
by: Mishkin, Aaron, et al.
Published: (2024)
Occam Gradient Descent
by: Kausik, B. N.
Published: (2024)
by: Kausik, B. N.
Published: (2024)
The Complexity Dynamics of Grokking
by: DeMoss, Branton, et al.
Published: (2024)
by: DeMoss, Branton, et al.
Published: (2024)
Measuring Sharpness in Grokking
by: Miller, Jack, et al.
Published: (2024)
by: Miller, Jack, et al.
Published: (2024)
Bridging Lottery Ticket and Grokking: Understanding Grokking from Inner Structure of Networks
by: Minegishi, Gouki, et al.
Published: (2023)
by: Minegishi, Gouki, et al.
Published: (2023)
Accelerating Natural Gradient Descent for PINNs with Randomized Numerical Linear Algebra
by: Bioli, Ivan, et al.
Published: (2025)
by: Bioli, Ivan, et al.
Published: (2025)
Accelerating Feedback-based Algorithms for Quantum Optimization Using Gradient Descent
by: Mozakka, Masih, et al.
Published: (2026)
by: Mozakka, Masih, et al.
Published: (2026)
Similar Items
-
WERank: Towards Rank Degradation Prevention for Self-Supervised Learning Using Weight Regularization
by: Pasand, Ali Saheb, et al.
Published: (2024) -
auto-fpt: Automating Free Probability Theory Calculations for Machine Learning Theory
by: Subramonian, Arjun, et al.
Published: (2025) -
Scalable Graph Self-Supervised Learning
by: Pasand, Ali Saheb, et al.
Published: (2024) -
Why Less is More (Sometimes): A Theory of Data Curation
by: Dohmatob, Elvis, et al.
Published: (2025) -
Model Collapse Demystified: The Case of Regression
by: Dohmatob, Elvis, et al.
Published: (2024)