Unified View of Grokking, Double Descent and Emergent Abilities: A Perspective from Circuits Competition
Fuente:
arXiv
Saved in:
| Main Authors: | Huang, Yufei, Hu, Shengding, Han, Xu, Liu, Zhiyuan, Sun, Maosong |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Stuffed Mamba: Oversized States Lead to the Inability to Forget
by: Chen, Yingfa, et al.
Published: (2024)
by: Chen, Yingfa, et al.
Published: (2024)
Two Speeds of Learning: A Representation-Readout Decomposition of Grokking and Double Descent
by: Chou, Chi-Ning, et al.
Published: (2026)
by: Chou, Chi-Ning, et al.
Published: (2026)
A Multi-Power Law for Loss Curve Prediction Across Learning Rate Schedules
by: Luo, Kairong, et al.
Published: (2025)
by: Luo, Kairong, et al.
Published: (2025)
Class-wise Activation Unravelling the Engima of Deep Double Descent
by: Gu, Yufei
Published: (2024)
by: Gu, Yufei
Published: (2024)
Egalitarian Gradient Descent: A Simple Approach to Accelerated Grokking
by: Pasand, Ali Saheb, et al.
Published: (2025)
by: Pasand, Ali Saheb, et al.
Published: (2025)
Predicting Emergent Abilities with Infinite Resolution Evaluation
by: Hu, Shengding, et al.
Published: (2023)
by: Hu, Shengding, et al.
Published: (2023)
To Grok Grokking: Provable Grokking in Ridge Regression
by: Xu, Mingyue, et al.
Published: (2026)
by: Xu, Mingyue, et al.
Published: (2026)
Unraveling the Enigma of Double Descent: An In-depth Analysis through the Lens of Learned Feature Space
by: Gu, Yufei, et al.
Published: (2023)
by: Gu, Yufei, et al.
Published: (2023)
LEGENT: Open Platform for Embodied Agents
by: Cheng, Zhili, et al.
Published: (2024)
by: Cheng, Zhili, et al.
Published: (2024)
ProSparse: Introducing and Enhancing Intrinsic Activation Sparsity within Large Language Models
by: Song, Chenyang, et al.
Published: (2024)
by: Song, Chenyang, et al.
Published: (2024)
Information-Theoretic Progress Measures reveal Grokking is an Emergent Phase Transition
by: Clauw, Kenzo, et al.
Published: (2024)
by: Clauw, Kenzo, et al.
Published: (2024)
Understanding Emergent Abilities of Language Models from the Loss Perspective
by: Du, Zhengxiao, et al.
Published: (2024)
by: Du, Zhengxiao, et al.
Published: (2024)
States Hidden in Hidden States: LLMs Emerge Discrete State Representations Implicitly
by: Chen, Junhao, et al.
Published: (2024)
by: Chen, Junhao, et al.
Published: (2024)
Why Do You Grok? A Theoretical Analysis of Grokking Modular Addition
by: Mohamadi, Mohamad Amin, et al.
Published: (2024)
by: Mohamadi, Mohamad Amin, et al.
Published: (2024)
Oversmoothing Alleviation in Graph Neural Networks: A Survey and Unified View
by: Jin, Yufei, et al.
Published: (2024)
by: Jin, Yufei, et al.
Published: (2024)
A Basin-Selection Perspective on Grokking via Singular Learning Theory
by: Cullen, Ben, et al.
Published: (2026)
by: Cullen, Ben, et al.
Published: (2026)
Exploring Grokking: Experimental and Mechanistic Investigations
by: Qiye, Hu, et al.
Published: (2024)
by: Qiye, Hu, et al.
Published: (2024)
Bridging Lottery Ticket and Grokking: Understanding Grokking from Inner Structure of Networks
by: Minegishi, Gouki, et al.
Published: (2023)
by: Minegishi, Gouki, et al.
Published: (2023)
Dichotomy of Early and Late Phase Implicit Biases Can Provably Induce Grokking
by: Lyu, Kaifeng, et al.
Published: (2023)
by: Lyu, Kaifeng, et al.
Published: (2023)
Towards Empirical Interpretation of Internal Circuits and Properties in Grokked Transformers on Modular Polynomials
by: Furuta, Hiroki, et al.
Published: (2024)
by: Furuta, Hiroki, et al.
Published: (2024)
Understanding Grokking Through A Robustness Viewpoint
by: Tan, Zhiquan, et al.
Published: (2023)
by: Tan, Zhiquan, et al.
Published: (2023)
Bayesian Double Descent
by: Polson, Nick, et al.
Published: (2025)
by: Polson, Nick, et al.
Published: (2025)
The Complexity Dynamics of Grokking
by: DeMoss, Branton, et al.
Published: (2024)
by: DeMoss, Branton, et al.
Published: (2024)
Measuring Sharpness in Grokking
by: Miller, Jack, et al.
Published: (2024)
by: Miller, Jack, et al.
Published: (2024)
Manipulating Sparse Double Descent
by: Zhang, Ya Shi
Published: (2024)
by: Zhang, Ya Shi
Published: (2024)
Let Me Grok for You: Accelerating Grokking via Embedding Transfer from a Weaker Model
by: Xu, Zhiwei, et al.
Published: (2025)
by: Xu, Zhiwei, et al.
Published: (2025)
Beyond Progress Measures: Theoretical Insights into the Mechanism of Grokking
by: Gu, Zihan, et al.
Published: (2025)
by: Gu, Zihan, et al.
Published: (2025)
A Bootstrap Perspective on Stochastic Gradient Descent
by: Lan, Hongjian, et al.
Published: (2025)
by: Lan, Hongjian, et al.
Published: (2025)
BurstAttention: An Efficient Distributed Attention Framework for Extremely Long Sequences
by: Sun, Ao, et al.
Published: (2024)
by: Sun, Ao, et al.
Published: (2024)
StateX: Enhancing RNN Recall via Post-training State Expansion
by: Shen, Xingyu, et al.
Published: (2025)
by: Shen, Xingyu, et al.
Published: (2025)
Grokked Models are Better Unlearners
by: Liang, Yuanbang, et al.
Published: (2025)
by: Liang, Yuanbang, et al.
Published: (2025)
A Unified Perspective on Optimization in Machine Learning and Neuroscience: From Gradient Descent to Neural Adaptation
by: Fernández, Jesús García, et al.
Published: (2025)
by: Fernández, Jesús García, et al.
Published: (2025)
Critical Data Size of Language Models from a Grokking Perspective
by: Zhu, Xuekai, et al.
Published: (2024)
by: Zhu, Xuekai, et al.
Published: (2024)
A Unified Approach to Controlling Implicit Regularization via Mirror Descent
by: Sun, Haoyuan, et al.
Published: (2023)
by: Sun, Haoyuan, et al.
Published: (2023)
Flatness is Necessary, Neural Collapse is Not: Rethinking Generalization via Grokking
by: Han, Ting, et al.
Published: (2025)
by: Han, Ting, et al.
Published: (2025)
Low-Dimensional and Transversely Curved Optimization Dynamics in Grokking
by: Xu, Yongzhong
Published: (2026)
by: Xu, Yongzhong
Published: (2026)
Topological Signatures of Grokking
by: Tang, Yifan, et al.
Published: (2026)
by: Tang, Yifan, et al.
Published: (2026)
Dropout Drops Double Descent
by: Yang, Tian-Le, et al.
Published: (2023)
by: Yang, Tian-Le, et al.
Published: (2023)
A Mirror Descent Perspective of Smoothed Sign Descent
by: Wang, Shuyang, et al.
Published: (2024)
by: Wang, Shuyang, et al.
Published: (2024)
ILDR: Geometric Early Detection of Grokking
by: Golwala, Shreel
Published: (2026)
by: Golwala, Shreel
Published: (2026)
Similar Items
-
Stuffed Mamba: Oversized States Lead to the Inability to Forget
by: Chen, Yingfa, et al.
Published: (2024) -
Two Speeds of Learning: A Representation-Readout Decomposition of Grokking and Double Descent
by: Chou, Chi-Ning, et al.
Published: (2026) -
A Multi-Power Law for Loss Curve Prediction Across Learning Rate Schedules
by: Luo, Kairong, et al.
Published: (2025) -
Class-wise Activation Unravelling the Engima of Deep Double Descent
by: Gu, Yufei
Published: (2024) -
Egalitarian Gradient Descent: A Simple Approach to Accelerated Grokking
by: Pasand, Ali Saheb, et al.
Published: (2025)