A Pre-Training Analogue of Grokking in Language Models: Tracing Delayed Grammatical Generalization
Fuente:
arXiv
Saved in:
| Main Authors: | Muckatira, Sherin, Shivagunde, Namrata, Deshpande, Vijeta, Rumshisky, Anna |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Beyond Perplexity: A Geometric and Spectral Study of Low-Rank Pre-Training
by: Shivagunde, Namrata, et al.
Published: (2026)
by: Shivagunde, Namrata, et al.
Published: (2026)
Emergent Abilities in Reduced-Scale Generative Language Models
by: Muckatira, Sherin, et al.
Published: (2024)
by: Muckatira, Sherin, et al.
Published: (2024)
Playing with Words, Improving with Rewards: Training Language Models for Creative Association
by: Deshpande, Vijeta, et al.
Published: (2026)
by: Deshpande, Vijeta, et al.
Published: (2026)
Deconstructing In-Context Learning: Understanding Prompts via Corruption
by: Shivagunde, Namrata, et al.
Published: (2024)
by: Shivagunde, Namrata, et al.
Published: (2024)
Analysis of student understanding in short‐answer explanations to concept questions using a human‐centered AI approach
by: Harpreet Auby, et al.
Published: (2025)
by: Harpreet Auby, et al.
Published: (2025)
Activation Steering for Synthetic Data Generation: The Role of Diversity in Downstream Safety Detection
by: Deshpande, Vijeta, et al.
Published: (2026)
by: Deshpande, Vijeta, et al.
Published: (2026)
Scaling Down to Scale Up: A Guide to Parameter-Efficient Fine-Tuning
by: Lialin, Vladislav, et al.
Published: (2023)
by: Lialin, Vladislav, et al.
Published: (2023)
Diverse, not Short: A Length-Controlled Data Selection Strategy for Improving Response Diversity of Language Models
by: Deshpande, Vijeta, et al.
Published: (2025)
by: Deshpande, Vijeta, et al.
Published: (2025)
The Norm-Separation Delay Law of Grokking: A First-Principles Theory of Delayed Generalization
by: Khanh, Truong Xuan, et al.
Published: (2026)
by: Khanh, Truong Xuan, et al.
Published: (2026)
Tracing the Path to Grokking: Embeddings, Dropout, and Network Activation
by: Salah, Ahmed, et al.
Published: (2025)
by: Salah, Ahmed, et al.
Published: (2025)
To Grok Grokking: Provable Grokking in Ridge Regression
by: Xu, Mingyue, et al.
Published: (2026)
by: Xu, Mingyue, et al.
Published: (2026)
Extra-Merge: Tracing the Rank-1 Subspace of Model Merging in Language Model Pre-Training
by: Zhou, Wenjie, et al.
Published: (2026)
by: Zhou, Wenjie, et al.
Published: (2026)
Grokked Models are Better Unlearners
by: Liang, Yuanbang, et al.
Published: (2025)
by: Liang, Yuanbang, et al.
Published: (2025)
A Penalty Goes a Long Way: Measuring Lexical Diversity in Synthetic Texts Under Prompt-Influenced Length Variations
by: Deshpande, Vijeta, et al.
Published: (2025)
by: Deshpande, Vijeta, et al.
Published: (2025)
The Complexity Dynamics of Grokking
by: DeMoss, Branton, et al.
Published: (2024)
by: DeMoss, Branton, et al.
Published: (2024)
Measuring Sharpness in Grokking
by: Miller, Jack, et al.
Published: (2024)
by: Miller, Jack, et al.
Published: (2024)
Bridging Lottery Ticket and Grokking: Understanding Grokking from Inner Structure of Networks
by: Minegishi, Gouki, et al.
Published: (2023)
by: Minegishi, Gouki, et al.
Published: (2023)
Prompt Perturbation Consistency Learning for Robust Language Models
by: Qiang, Yao, et al.
Published: (2024)
by: Qiang, Yao, et al.
Published: (2024)
LocalTweets to LocalHealth: A Mental Health Surveillance Framework Based on Twitter Data
by: Deshpande, Vijeta, et al.
Published: (2024)
by: Deshpande, Vijeta, et al.
Published: (2024)
Grokking in Linear Models for Logistic Regression
by: Das, Nataraj, et al.
Published: (2026)
by: Das, Nataraj, et al.
Published: (2026)
Grokking as the Transition from Lazy to Rich Training Dynamics
by: Kumar, Tanishq, et al.
Published: (2023)
by: Kumar, Tanishq, et al.
Published: (2023)
First-Passage Prediction of Grokking Delay: ACalibrated Law under AdamW with Causal Validation
by: Khanh, Truong Xuan, et al.
Published: (2026)
by: Khanh, Truong Xuan, et al.
Published: (2026)
Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks
by: Shandilya, Utkarsh, et al.
Published: (2025)
by: Shandilya, Utkarsh, et al.
Published: (2025)
Flatness is Necessary, Neural Collapse is Not: Rethinking Generalization via Grokking
by: Han, Ting, et al.
Published: (2025)
by: Han, Ting, et al.
Published: (2025)
Deep Grokking: Would Deep Neural Networks Generalize Better?
by: Fan, Simin, et al.
Published: (2024)
by: Fan, Simin, et al.
Published: (2024)
Grokking in LLM Pretraining? Monitor Memorization-to-Generalization without Test
by: Li, Ziyue, et al.
Published: (2025)
by: Li, Ziyue, et al.
Published: (2025)
Grokking of Diffusion Models: Case Study on Modular Addition
by: Kim, Joon Hyeok, et al.
Published: (2026)
by: Kim, Joon Hyeok, et al.
Published: (2026)
Topological Signatures of Grokking
by: Tang, Yifan, et al.
Published: (2026)
by: Tang, Yifan, et al.
Published: (2026)
Late-Stage Generalization Collapse in Grokking: Detecting anti-grokking with Weightwatcher
by: Prakash, Hari K, et al.
Published: (2026)
by: Prakash, Hari K, et al.
Published: (2026)
ILDR: Geometric Early Detection of Grokking
by: Golwala, Shreel
Published: (2026)
by: Golwala, Shreel
Published: (2026)
Exploring Grokking: Experimental and Mechanistic Investigations
by: Qiye, Hu, et al.
Published: (2024)
by: Qiye, Hu, et al.
Published: (2024)
Grokking Beyond Neural Networks: An Empirical Exploration with Model Complexity
by: Miller, Jack, et al.
Published: (2023)
by: Miller, Jack, et al.
Published: (2023)
Critical Data Size of Language Models from a Grokking Perspective
by: Zhu, Xuekai, et al.
Published: (2024)
by: Zhu, Xuekai, et al.
Published: (2024)
Grokking Explained: A Statistical Phenomenon
by: Carvalho, Breno W., et al.
Published: (2025)
by: Carvalho, Breno W., et al.
Published: (2025)
Global Vegetation Modeling with Pre-Trained Weather Transformers
by: Janetzky, Pascal, et al.
Published: (2024)
by: Janetzky, Pascal, et al.
Published: (2024)
Grokking and Generalization Collapse: Insights from \texttt{HTSR} theory
by: Prakash, Hari K., et al.
Published: (2025)
by: Prakash, Hari K., et al.
Published: (2025)
Model Capacity Determines Grokking through Competing Memorisation and Generalisation Speeds
by: Song, Yiding, et al.
Published: (2026)
by: Song, Yiding, et al.
Published: (2026)
DelayPTC-LLM: Metro Passenger Travel Choice Prediction under Train Delays with Large Language Models
by: Chen, Chen, et al.
Published: (2024)
by: Chen, Chen, et al.
Published: (2024)
Distributional Spectral Diagnostics for Localizing Grokking Transitions
by: Wang, Ziyue, et al.
Published: (2026)
by: Wang, Ziyue, et al.
Published: (2026)
GrokAlign: Geometric Characterisation and Acceleration of Grokking
by: Walker, Thomas, et al.
Published: (2025)
by: Walker, Thomas, et al.
Published: (2025)
Similar Items
-
Beyond Perplexity: A Geometric and Spectral Study of Low-Rank Pre-Training
by: Shivagunde, Namrata, et al.
Published: (2026) -
Emergent Abilities in Reduced-Scale Generative Language Models
by: Muckatira, Sherin, et al.
Published: (2024) -
Playing with Words, Improving with Rewards: Training Language Models for Creative Association
by: Deshpande, Vijeta, et al.
Published: (2026) -
Deconstructing In-Context Learning: Understanding Prompts via Corruption
by: Shivagunde, Namrata, et al.
Published: (2024) -
Analysis of student understanding in short‐answer explanations to concept questions using a human‐centered AI approach
by: Harpreet Auby, et al.
Published: (2025)