Gespeichert in:
| Hauptverfasser: | Miller, Jack, Gleeson, Patrick, O'Neill, Charles, Bui, Thang, Levi, Noam |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2402.08946 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Grokking Beyond Neural Networks: An Empirical Exploration with Model Complexity
von: Miller, Jack, et al.
Veröffentlicht: (2023)
von: Miller, Jack, et al.
Veröffentlicht: (2023)
Sparse Autoencoders Enable Scalable and Reliable Circuit Identification in Language Models
von: O'Neill, Charles, et al.
Veröffentlicht: (2024)
von: O'Neill, Charles, et al.
Veröffentlicht: (2024)
Grokking at the Edge of Linear Separability
von: Beck, Alon, et al.
Veröffentlicht: (2024)
von: Beck, Alon, et al.
Veröffentlicht: (2024)
Self-Attention as a Parametric Endofunctor: A Categorical Framework for Transformer Architectures
von: O'Neill, Charles
Veröffentlicht: (2025)
von: O'Neill, Charles
Veröffentlicht: (2025)
Grokking in Linear Estimators -- A Solvable Model that Groks without Understanding
von: Levi, Noam, et al.
Veröffentlicht: (2023)
von: Levi, Noam, et al.
Veröffentlicht: (2023)
Type 2 Tobit Sample Selection Models with Bayesian Additive Regression Trees
von: O'Neill, Eoghan
Veröffentlicht: (2025)
von: O'Neill, Eoghan
Veröffentlicht: (2025)
Compute Optimal Inference and Provable Amortisation Gap in Sparse Autoencoders
von: O'Neill, Charles, et al.
Veröffentlicht: (2024)
von: O'Neill, Charles, et al.
Veröffentlicht: (2024)
Resurrecting the Salmon: Rethinking Mechanistic Interpretability with Domain-Specific Sparse Autoencoders
von: O'Neill, Charles, et al.
Veröffentlicht: (2025)
von: O'Neill, Charles, et al.
Veröffentlicht: (2025)
Likelihood approximations via Gaussian approximate inference
von: Bui, Thang D.
Veröffentlicht: (2024)
von: Bui, Thang D.
Veröffentlicht: (2024)
A Simple Model of Inference Scaling Laws
von: Levi, Noam
Veröffentlicht: (2024)
von: Levi, Noam
Veröffentlicht: (2024)
Learning Shrinks the Hard Tail: Training-Dependent Inference Scaling in a Solvable Linear Model
von: Levi, Noam
Veröffentlicht: (2026)
von: Levi, Noam
Veröffentlicht: (2026)
Disentangling Dense Embeddings with Sparse Autoencoders
von: O'Neill, Charles, et al.
Veröffentlicht: (2024)
von: O'Neill, Charles, et al.
Veröffentlicht: (2024)
Sketching the Heat Kernel: Using Gaussian Processes to Embed Data
von: Gilbert, Anna C., et al.
Veröffentlicht: (2024)
von: Gilbert, Anna C., et al.
Veröffentlicht: (2024)
Beyond ReinMax: Low-Variance Gradient Estimators for Discrete Latent Variables
von: Wang, Daniel, et al.
Veröffentlicht: (2026)
von: Wang, Daniel, et al.
Veröffentlicht: (2026)
Modelling the Doughnut of social and planetary boundaries with frugal machine learning
von: Vrizzi, Stefano, et al.
Veröffentlicht: (2025)
von: Vrizzi, Stefano, et al.
Veröffentlicht: (2025)
From superposition to sparse codes: interpretable representations in neural networks
von: Klindt, David, et al.
Veröffentlicht: (2025)
von: Klindt, David, et al.
Veröffentlicht: (2025)
To Grok Grokking: Provable Grokking in Ridge Regression
von: Xu, Mingyue, et al.
Veröffentlicht: (2026)
von: Xu, Mingyue, et al.
Veröffentlicht: (2026)
Beyond Progress Measures: Theoretical Insights into the Mechanism of Grokking
von: Gu, Zihan, et al.
Veröffentlicht: (2025)
von: Gu, Zihan, et al.
Veröffentlicht: (2025)
Progress Measures for Grokking on Real-world Tasks
von: Golechha, Satvik
Veröffentlicht: (2024)
von: Golechha, Satvik
Veröffentlicht: (2024)
A Single Direction of Truth: An Observer Model's Linear Residual Probe Exposes and Steers Contextual Hallucinations
von: O'Neill, Charles, et al.
Veröffentlicht: (2025)
von: O'Neill, Charles, et al.
Veröffentlicht: (2025)
Sparse Gaussian Processes: Structured Approximations and Power-EP Revisited
von: Bui, Thang D., et al.
Veröffentlicht: (2025)
von: Bui, Thang D., et al.
Veröffentlicht: (2025)
CA-PCA: Manifold Dimension Estimation, Adapted for Curvature
von: Gilbert, Anna C., et al.
Veröffentlicht: (2023)
von: Gilbert, Anna C., et al.
Veröffentlicht: (2023)
Low-Rank Key Value Attention
von: O'Neill, James, et al.
Veröffentlicht: (2026)
von: O'Neill, James, et al.
Veröffentlicht: (2026)
The Complexity Dynamics of Grokking
von: DeMoss, Branton, et al.
Veröffentlicht: (2024)
von: DeMoss, Branton, et al.
Veröffentlicht: (2024)
Bridging Lottery Ticket and Grokking: Understanding Grokking from Inner Structure of Networks
von: Minegishi, Gouki, et al.
Veröffentlicht: (2023)
von: Minegishi, Gouki, et al.
Veröffentlicht: (2023)
The Implicit Bias of Logit Regularization
von: Beck, Alon, et al.
Veröffentlicht: (2026)
von: Beck, Alon, et al.
Veröffentlicht: (2026)
Classifying Overlapping Gaussian Mixtures in High Dimensions: From Optimal Classifiers to Neural Nets
von: Cohen, Khen, et al.
Veröffentlicht: (2024)
von: Cohen, Khen, et al.
Veröffentlicht: (2024)
Late-Stage Generalization Collapse in Grokking: Detecting anti-grokking with Weightwatcher
von: Prakash, Hari K, et al.
Veröffentlicht: (2026)
von: Prakash, Hari K, et al.
Veröffentlicht: (2026)
The Underlying Scaling Laws and Universal Statistical Structure of Complex Datasets
von: Levi, Noam, et al.
Veröffentlicht: (2023)
von: Levi, Noam, et al.
Veröffentlicht: (2023)
Grokking and Generalization Collapse: Insights from \texttt{HTSR} theory
von: Prakash, Hari K., et al.
Veröffentlicht: (2025)
von: Prakash, Hari K., et al.
Veröffentlicht: (2025)
Tighter sparse variational Gaussian processes
von: Bui, Thang D., et al.
Veröffentlicht: (2025)
von: Bui, Thang D., et al.
Veröffentlicht: (2025)
Grokked Models are Better Unlearners
von: Liang, Yuanbang, et al.
Veröffentlicht: (2025)
von: Liang, Yuanbang, et al.
Veröffentlicht: (2025)
Topological Signatures of Grokking
von: Tang, Yifan, et al.
Veröffentlicht: (2026)
von: Tang, Yifan, et al.
Veröffentlicht: (2026)
Information-Theoretic Progress Measures reveal Grokking is an Emergent Phase Transition
von: Clauw, Kenzo, et al.
Veröffentlicht: (2024)
von: Clauw, Kenzo, et al.
Veröffentlicht: (2024)
Pretraining Scaling Laws for Generative Evaluations of Language Models
von: Schaeffer, Rylan, et al.
Veröffentlicht: (2025)
von: Schaeffer, Rylan, et al.
Veröffentlicht: (2025)
Decoupled Weight Decay for Any $p$ Norm
von: Outmezguine, Nadav Joseph, et al.
Veröffentlicht: (2024)
von: Outmezguine, Nadav Joseph, et al.
Veröffentlicht: (2024)
Exploring Grokking: Experimental and Mechanistic Investigations
von: Qiye, Hu, et al.
Veröffentlicht: (2024)
von: Qiye, Hu, et al.
Veröffentlicht: (2024)
ILDR: Geometric Early Detection of Grokking
von: Golwala, Shreel
Veröffentlicht: (2026)
von: Golwala, Shreel
Veröffentlicht: (2026)
Distributional Spectral Diagnostics for Localizing Grokking Transitions
von: Wang, Ziyue, et al.
Veröffentlicht: (2026)
von: Wang, Ziyue, et al.
Veröffentlicht: (2026)
GrokAlign: Geometric Characterisation and Acceleration of Grokking
von: Walker, Thomas, et al.
Veröffentlicht: (2025)
von: Walker, Thomas, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Grokking Beyond Neural Networks: An Empirical Exploration with Model Complexity
von: Miller, Jack, et al.
Veröffentlicht: (2023) -
Sparse Autoencoders Enable Scalable and Reliable Circuit Identification in Language Models
von: O'Neill, Charles, et al.
Veröffentlicht: (2024) -
Grokking at the Edge of Linear Separability
von: Beck, Alon, et al.
Veröffentlicht: (2024) -
Self-Attention as a Parametric Endofunctor: A Categorical Framework for Transformer Architectures
von: O'Neill, Charles
Veröffentlicht: (2025) -
Grokking in Linear Estimators -- A Solvable Model that Groks without Understanding
von: Levi, Noam, et al.
Veröffentlicht: (2023)