Explaining Grokking in Transformers through the Lens of Inductive Bias
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Singh, Jaisidh, Misra, Diganta, Orvieto, Antonio |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
(Almost) Free Modality Stitching of Foundation Models
von: Singh, Jaisidh, et al.
Veröffentlicht: (2025)
von: Singh, Jaisidh, et al.
Veröffentlicht: (2025)
On the low-shot transferability of [V]-Mamba
von: Misra, Diganta, et al.
Veröffentlicht: (2024)
von: Misra, Diganta, et al.
Veröffentlicht: (2024)
GRASP: Deterministic argument ranking in interaction graphs
von: Misra, Diganta, et al.
Veröffentlicht: (2026)
von: Misra, Diganta, et al.
Veröffentlicht: (2026)
Adam Simplified: Bias Correction Debunked
von: Laing, Sam, et al.
Veröffentlicht: (2025)
von: Laing, Sam, et al.
Veröffentlicht: (2025)
Explaining Grokking and Information Bottleneck through Neural Collapse Emergence
von: Sakamoto, Keitaro, et al.
Veröffentlicht: (2025)
von: Sakamoto, Keitaro, et al.
Veröffentlicht: (2025)
Geometric Inductive Biases of Deep Networks: The Role of Data and Architecture
von: Movahedi, Sajad, et al.
Veröffentlicht: (2024)
von: Movahedi, Sajad, et al.
Veröffentlicht: (2024)
The Geometric Inductive Bias of Grokking: Bypassing Phase Transitions via Architectural Topology
von: Yıldırım, Alper
Veröffentlicht: (2026)
von: Yıldırım, Alper
Veröffentlicht: (2026)
Grokking Explained: A Statistical Phenomenon
von: Carvalho, Breno W., et al.
Veröffentlicht: (2025)
von: Carvalho, Breno W., et al.
Veröffentlicht: (2025)
Adaptive Methods through the Lens of SDEs: Theoretical Insights on the Role of Noise
von: Compagnoni, Enea Monzio, et al.
Veröffentlicht: (2024)
von: Compagnoni, Enea Monzio, et al.
Veröffentlicht: (2024)
Universal Dynamics of Warmup Stable Decay: understanding WSD beyond Transformers
von: Belloni, Annalisa, et al.
Veröffentlicht: (2026)
von: Belloni, Annalisa, et al.
Veröffentlicht: (2026)
Revisiting associative recall in modern recurrent models
von: Okpekpe, Destiny, et al.
Veröffentlicht: (2025)
von: Okpekpe, Destiny, et al.
Veröffentlicht: (2025)
Generalized Linear Mode Connectivity for Transformers
von: Theus, Alexander, et al.
Veröffentlicht: (2025)
von: Theus, Alexander, et al.
Veröffentlicht: (2025)
To Grok Grokking: Provable Grokking in Ridge Regression
von: Xu, Mingyue, et al.
Veröffentlicht: (2026)
von: Xu, Mingyue, et al.
Veröffentlicht: (2026)
MechPert: Mechanistic Consensus as an Inductive Bias for Unseen Perturbation Prediction
von: Martell, Marc Boubnovski, et al.
Veröffentlicht: (2026)
von: Martell, Marc Boubnovski, et al.
Veröffentlicht: (2026)
An Adaptive Stochastic Gradient Method with Non-negative Gauss-Newton Stepsizes
von: Orvieto, Antonio, et al.
Veröffentlicht: (2024)
von: Orvieto, Antonio, et al.
Veröffentlicht: (2024)
In Search of Adam's Secret Sauce
von: Orvieto, Antonio, et al.
Veröffentlicht: (2025)
von: Orvieto, Antonio, et al.
Veröffentlicht: (2025)
Dataset Difficulty and the Role of Inductive Bias
von: Kwok, Devin, et al.
Veröffentlicht: (2024)
von: Kwok, Devin, et al.
Veröffentlicht: (2024)
Interpolated-MLPs: Controllable Inductive Bias
von: Wu, Sean, et al.
Veröffentlicht: (2024)
von: Wu, Sean, et al.
Veröffentlicht: (2024)
Towards Exact Computation of Inductive Bias
von: Boopathy, Akhilan, et al.
Veröffentlicht: (2024)
von: Boopathy, Akhilan, et al.
Veröffentlicht: (2024)
Compositional Sparsity as an Inductive Bias for Neural Architecture Design
von: Lin, Hongyu, et al.
Veröffentlicht: (2026)
von: Lin, Hongyu, et al.
Veröffentlicht: (2026)
When Data Falls Short: Grokking Below the Critical Threshold
von: Singh, Vaibhav, et al.
Veröffentlicht: (2025)
von: Singh, Vaibhav, et al.
Veröffentlicht: (2025)
Recurrent neural networks: vanishing and exploding gradients are not the end of the story
von: Zucchet, Nicolas, et al.
Veröffentlicht: (2024)
von: Zucchet, Nicolas, et al.
Veröffentlicht: (2024)
Improved state mixing in higher-order and block diagonal linear recurrent networks
von: Dubinin, Igor, et al.
Veröffentlicht: (2026)
von: Dubinin, Igor, et al.
Veröffentlicht: (2026)
An Uncertainty Principle for Linear Recurrent Neural Networks
von: François, Alexandre, et al.
Veröffentlicht: (2025)
von: François, Alexandre, et al.
Veröffentlicht: (2025)
When, Where and Why to Average Weights?
von: Ajroldi, Niccolò, et al.
Veröffentlicht: (2025)
von: Ajroldi, Niccolò, et al.
Veröffentlicht: (2025)
Theoretical Investigation on Inductive Bias of Isolation Forest
von: Zheng, Qin-Cheng, et al.
Veröffentlicht: (2025)
von: Zheng, Qin-Cheng, et al.
Veröffentlicht: (2025)
NeuralGrok: Accelerate Grokking by Neural Gradient Transformation
von: Zhou, Xinyu, et al.
Veröffentlicht: (2025)
von: Zhou, Xinyu, et al.
Veröffentlicht: (2025)
The Complexity Dynamics of Grokking
von: DeMoss, Branton, et al.
Veröffentlicht: (2024)
von: DeMoss, Branton, et al.
Veröffentlicht: (2024)
Measuring Sharpness in Grokking
von: Miller, Jack, et al.
Veröffentlicht: (2024)
von: Miller, Jack, et al.
Veröffentlicht: (2024)
Bridging Lottery Ticket and Grokking: Understanding Grokking from Inner Structure of Networks
von: Minegishi, Gouki, et al.
Veröffentlicht: (2023)
von: Minegishi, Gouki, et al.
Veröffentlicht: (2023)
Algorithmic Task Capture, Computational Complexity, and Inductive Bias of Infinite Transformers
von: Davidovich, Orit, et al.
Veröffentlicht: (2026)
von: Davidovich, Orit, et al.
Veröffentlicht: (2026)
Energy-Gated Attention: Spectral Salience as an Inductive Bias for Transformer Attention
von: Zeris, Athanasios
Veröffentlicht: (2026)
von: Zeris, Athanasios
Veröffentlicht: (2026)
Towards Understanding Inductive Bias in Transformers: A View From Infinity
von: Lavie, Itay, et al.
Veröffentlicht: (2024)
von: Lavie, Itay, et al.
Veröffentlicht: (2024)
Shaping Inductive Bias in Diffusion Models through Frequency-Based Noise Control
von: Jiralerspong, Thomas, et al.
Veröffentlicht: (2025)
von: Jiralerspong, Thomas, et al.
Veröffentlicht: (2025)
Model Capacity Determines Grokking through Competing Memorisation and Generalisation Speeds
von: Song, Yiding, et al.
Veröffentlicht: (2026)
von: Song, Yiding, et al.
Veröffentlicht: (2026)
Agents Learn Their Runtime: Interpreter Persistence as Training-Time Semantics
von: May, Victor, et al.
Veröffentlicht: (2026)
von: May, Victor, et al.
Veröffentlicht: (2026)
Data Distributional Properties As Inductive Bias for Systematic Generalization
von: del Rio, Felipe, et al.
Veröffentlicht: (2025)
von: del Rio, Felipe, et al.
Veröffentlicht: (2025)
Graph Classification with GNNs: Optimisation, Representation and Inductive Bias
von: a, P. Krishna Kumar, et al.
Veröffentlicht: (2024)
von: a, P. Krishna Kumar, et al.
Veröffentlicht: (2024)
Grokking as Structural Inference: Transformers Need Bayesian Lottery Tickets
von: Hidajat, Kai, et al.
Veröffentlicht: (2026)
von: Hidajat, Kai, et al.
Veröffentlicht: (2026)
Grokked Models are Better Unlearners
von: Liang, Yuanbang, et al.
Veröffentlicht: (2025)
von: Liang, Yuanbang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
(Almost) Free Modality Stitching of Foundation Models
von: Singh, Jaisidh, et al.
Veröffentlicht: (2025) -
On the low-shot transferability of [V]-Mamba
von: Misra, Diganta, et al.
Veröffentlicht: (2024) -
GRASP: Deterministic argument ranking in interaction graphs
von: Misra, Diganta, et al.
Veröffentlicht: (2026) -
Adam Simplified: Bias Correction Debunked
von: Laing, Sam, et al.
Veröffentlicht: (2025) -
Explaining Grokking and Information Bottleneck through Neural Collapse Emergence
von: Sakamoto, Keitaro, et al.
Veröffentlicht: (2025)