Grokking Beyond Neural Networks: An Empirical Exploration with Model Complexity
Fuente:
arXiv
Saved in:
| Main Authors: | Miller, Jack, O'Neill, Charles, Bui, Thang |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Measuring Sharpness in Grokking
by: Miller, Jack, et al.
Published: (2024)
by: Miller, Jack, et al.
Published: (2024)
Sparse Autoencoders Enable Scalable and Reliable Circuit Identification in Language Models
by: O'Neill, Charles, et al.
Published: (2024)
by: O'Neill, Charles, et al.
Published: (2024)
Self-Attention as a Parametric Endofunctor: A Categorical Framework for Transformer Architectures
by: O'Neill, Charles
Published: (2025)
by: O'Neill, Charles
Published: (2025)
Type 2 Tobit Sample Selection Models with Bayesian Additive Regression Trees
by: O'Neill, Eoghan
Published: (2025)
by: O'Neill, Eoghan
Published: (2025)
Beyond ReinMax: Low-Variance Gradient Estimators for Discrete Latent Variables
by: Wang, Daniel, et al.
Published: (2026)
by: Wang, Daniel, et al.
Published: (2026)
The Complexity Dynamics of Grokking
by: DeMoss, Branton, et al.
Published: (2024)
by: DeMoss, Branton, et al.
Published: (2024)
Resurrecting the Salmon: Rethinking Mechanistic Interpretability with Domain-Specific Sparse Autoencoders
by: O'Neill, Charles, et al.
Published: (2025)
by: O'Neill, Charles, et al.
Published: (2025)
Compute Optimal Inference and Provable Amortisation Gap in Sparse Autoencoders
by: O'Neill, Charles, et al.
Published: (2024)
by: O'Neill, Charles, et al.
Published: (2024)
Likelihood approximations via Gaussian approximate inference
by: Bui, Thang D.
Published: (2024)
by: Bui, Thang D.
Published: (2024)
Deep Grokking: Would Deep Neural Networks Generalize Better?
by: Fan, Simin, et al.
Published: (2024)
by: Fan, Simin, et al.
Published: (2024)
Modelling the Doughnut of social and planetary boundaries with frugal machine learning
by: Vrizzi, Stefano, et al.
Published: (2025)
by: Vrizzi, Stefano, et al.
Published: (2025)
Bridging Lottery Ticket and Grokking: Understanding Grokking from Inner Structure of Networks
by: Minegishi, Gouki, et al.
Published: (2023)
by: Minegishi, Gouki, et al.
Published: (2023)
Feature Repulsion and Spectral Lock-in: An Empirical Study of Two-Layer Network Grokking
by: Xu, Yongzhong
Published: (2026)
by: Xu, Yongzhong
Published: (2026)
Grokking Beyond the Euclidean Norm of Model Parameters
by: Notsawo, Pascal Jr Tikeng, et al.
Published: (2025)
by: Notsawo, Pascal Jr Tikeng, et al.
Published: (2025)
Beyond Progress Measures: Theoretical Insights into the Mechanism of Grokking
by: Gu, Zihan, et al.
Published: (2025)
by: Gu, Zihan, et al.
Published: (2025)
Disentangling Dense Embeddings with Sparse Autoencoders
by: O'Neill, Charles, et al.
Published: (2024)
by: O'Neill, Charles, et al.
Published: (2024)
A Single Direction of Truth: An Observer Model's Linear Residual Probe Exposes and Steers Contextual Hallucinations
by: O'Neill, Charles, et al.
Published: (2025)
by: O'Neill, Charles, et al.
Published: (2025)
A Systematic Empirical Study of Grokking: Depth, Architecture, Activation, and Regularization
by: Manir, Shalima Binta, et al.
Published: (2026)
by: Manir, Shalima Binta, et al.
Published: (2026)
To Grok Grokking: Provable Grokking in Ridge Regression
by: Xu, Mingyue, et al.
Published: (2026)
by: Xu, Mingyue, et al.
Published: (2026)
Deep Learning Through A Telescoping Lens: A Simple Model Provides Empirical Insights On Grokking, Gradient Boosting & Beyond
by: Jeffares, Alan, et al.
Published: (2024)
by: Jeffares, Alan, et al.
Published: (2024)
Sketching the Heat Kernel: Using Gaussian Processes to Embed Data
by: Gilbert, Anna C., et al.
Published: (2024)
by: Gilbert, Anna C., et al.
Published: (2024)
Grokked Models are Better Unlearners
by: Liang, Yuanbang, et al.
Published: (2025)
by: Liang, Yuanbang, et al.
Published: (2025)
From superposition to sparse codes: interpretable representations in neural networks
by: Klindt, David, et al.
Published: (2025)
by: Klindt, David, et al.
Published: (2025)
NeuralGrok: Accelerate Grokking by Neural Gradient Transformation
by: Zhou, Xinyu, et al.
Published: (2025)
by: Zhou, Xinyu, et al.
Published: (2025)
Sparse Gaussian Processes: Structured Approximations and Power-EP Revisited
by: Bui, Thang D., et al.
Published: (2025)
by: Bui, Thang D., et al.
Published: (2025)
Explaining Grokking and Information Bottleneck through Neural Collapse Emergence
by: Sakamoto, Keitaro, et al.
Published: (2025)
by: Sakamoto, Keitaro, et al.
Published: (2025)
Flatness is Necessary, Neural Collapse is Not: Rethinking Generalization via Grokking
by: Han, Ting, et al.
Published: (2025)
by: Han, Ting, et al.
Published: (2025)
Grokking as Dimensional Phase Transition in Neural Networks
by: Wang, Ping
Published: (2026)
by: Wang, Ping
Published: (2026)
Towards Empirical Interpretation of Internal Circuits and Properties in Grokked Transformers on Modular Polynomials
by: Furuta, Hiroki, et al.
Published: (2024)
by: Furuta, Hiroki, et al.
Published: (2024)
Late-Stage Generalization Collapse in Grokking: Detecting anti-grokking with Weightwatcher
by: Prakash, Hari K, et al.
Published: (2026)
by: Prakash, Hari K, et al.
Published: (2026)
Low-Rank Key Value Attention
by: O'Neill, James, et al.
Published: (2026)
by: O'Neill, James, et al.
Published: (2026)
CA-PCA: Manifold Dimension Estimation, Adapted for Curvature
by: Gilbert, Anna C., et al.
Published: (2023)
by: Gilbert, Anna C., et al.
Published: (2023)
Grokking in Linear Models for Logistic Regression
by: Das, Nataraj, et al.
Published: (2026)
by: Das, Nataraj, et al.
Published: (2026)
Grokking and Generalization Collapse: Insights from \texttt{HTSR} theory
by: Prakash, Hari K., et al.
Published: (2025)
by: Prakash, Hari K., et al.
Published: (2025)
Grokking of Diffusion Models: Case Study on Modular Addition
by: Kim, Joon Hyeok, et al.
Published: (2026)
by: Kim, Joon Hyeok, et al.
Published: (2026)
Tracing the Path to Grokking: Embeddings, Dropout, and Network Activation
by: Salah, Ahmed, et al.
Published: (2025)
by: Salah, Ahmed, et al.
Published: (2025)
Topological Signatures of Grokking
by: Tang, Yifan, et al.
Published: (2026)
by: Tang, Yifan, et al.
Published: (2026)
Exploring Grokking: Experimental and Mechanistic Investigations
by: Qiye, Hu, et al.
Published: (2024)
by: Qiye, Hu, et al.
Published: (2024)
ILDR: Geometric Early Detection of Grokking
by: Golwala, Shreel
Published: (2026)
by: Golwala, Shreel
Published: (2026)
Tighter sparse variational Gaussian processes
by: Bui, Thang D., et al.
Published: (2025)
by: Bui, Thang D., et al.
Published: (2025)
Similar Items
-
Measuring Sharpness in Grokking
by: Miller, Jack, et al.
Published: (2024) -
Sparse Autoencoders Enable Scalable and Reliable Circuit Identification in Language Models
by: O'Neill, Charles, et al.
Published: (2024) -
Self-Attention as a Parametric Endofunctor: A Categorical Framework for Transformer Architectures
by: O'Neill, Charles
Published: (2025) -
Type 2 Tobit Sample Selection Models with Bayesian Additive Regression Trees
by: O'Neill, Eoghan
Published: (2025) -
Beyond ReinMax: Low-Variance Gradient Estimators for Discrete Latent Variables
by: Wang, Daniel, et al.
Published: (2026)