Grokking in LLM Pretraining? Monitor Memorization-to-Generalization without Test
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Ziyue, Fan, Chenrui, Zhou, Tianyi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Skip a Layer or Loop it? Test-Time Depth Adaptation of Pretrained LLMs
von: Li, Ziyue, et al.
Veröffentlicht: (2025)
von: Li, Ziyue, et al.
Veröffentlicht: (2025)
Your Mixture-of-Experts LLM Is Secretly an Embedding Model For Free
von: Li, Ziyue, et al.
Veröffentlicht: (2024)
von: Li, Ziyue, et al.
Veröffentlicht: (2024)
R2-T2: Re-Routing in Test-Time for Multimodal Mixture-of-Experts
von: Li, Zhongyang, et al.
Veröffentlicht: (2025)
von: Li, Zhongyang, et al.
Veröffentlicht: (2025)
Distributional Spectral Diagnostics for Localizing Grokking Transitions
von: Wang, Ziyue, et al.
Veröffentlicht: (2026)
von: Wang, Ziyue, et al.
Veröffentlicht: (2026)
C3PO: Critical-Layer, Core-Expert, Collaborative Pathway Optimization for Test-Time Expert Re-Mixing
von: Li, Zhongyang, et al.
Veröffentlicht: (2025)
von: Li, Zhongyang, et al.
Veröffentlicht: (2025)
Routing Manifold Alignment Improves Generalization of Mixture-of-Experts LLMs
von: Li, Zhongyang, et al.
Veröffentlicht: (2025)
von: Li, Zhongyang, et al.
Veröffentlicht: (2025)
Missing Premise exacerbates Overthinking: Are Reasoning Models losing Critical Thinking Skill?
von: Fan, Chenrui, et al.
Veröffentlicht: (2025)
von: Fan, Chenrui, et al.
Veröffentlicht: (2025)
Deep Grokking: Would Deep Neural Networks Generalize Better?
von: Fan, Simin, et al.
Veröffentlicht: (2024)
von: Fan, Simin, et al.
Veröffentlicht: (2024)
To Grok Grokking: Provable Grokking in Ridge Regression
von: Xu, Mingyue, et al.
Veröffentlicht: (2026)
von: Xu, Mingyue, et al.
Veröffentlicht: (2026)
Schoenfeld's Anatomy of Mathematical Reasoning by Language Models
von: Li, Ming, et al.
Veröffentlicht: (2025)
von: Li, Ming, et al.
Veröffentlicht: (2025)
NeuralGrok: Accelerate Grokking by Neural Gradient Transformation
von: Zhou, Xinyu, et al.
Veröffentlicht: (2025)
von: Zhou, Xinyu, et al.
Veröffentlicht: (2025)
Memorization Dynamics of Fill-in-the-Middle Pretraining
von: von Arx, Tobias, et al.
Veröffentlicht: (2026)
von: von Arx, Tobias, et al.
Veröffentlicht: (2026)
Grokked Models are Better Unlearners
von: Liang, Yuanbang, et al.
Veröffentlicht: (2025)
von: Liang, Yuanbang, et al.
Veröffentlicht: (2025)
Exploring Grokking: Experimental and Mechanistic Investigations
von: Qiye, Hu, et al.
Veröffentlicht: (2024)
von: Qiye, Hu, et al.
Veröffentlicht: (2024)
The Complexity Dynamics of Grokking
von: DeMoss, Branton, et al.
Veröffentlicht: (2024)
von: DeMoss, Branton, et al.
Veröffentlicht: (2024)
Measuring Sharpness in Grokking
von: Miller, Jack, et al.
Veröffentlicht: (2024)
von: Miller, Jack, et al.
Veröffentlicht: (2024)
Bridging Lottery Ticket and Grokking: Understanding Grokking from Inner Structure of Networks
von: Minegishi, Gouki, et al.
Veröffentlicht: (2023)
von: Minegishi, Gouki, et al.
Veröffentlicht: (2023)
Grokking Group Multiplication with Cosets
von: Stander, Dashiell, et al.
Veröffentlicht: (2023)
von: Stander, Dashiell, et al.
Veröffentlicht: (2023)
Memorization Sinks: Isolating Memorization during LLM Training
von: Ghosal, Gaurav R., et al.
Veröffentlicht: (2025)
von: Ghosal, Gaurav R., et al.
Veröffentlicht: (2025)
Generalization v.s. Memorization: Tracing Language Models' Capabilities Back to Pretraining Data
von: Wang, Xinyi, et al.
Veröffentlicht: (2024)
von: Wang, Xinyi, et al.
Veröffentlicht: (2024)
To Memorize or to Retrieve: Scaling Laws for RAG-Considerate Pretraining
von: Singh, Karan, et al.
Veröffentlicht: (2026)
von: Singh, Karan, et al.
Veröffentlicht: (2026)
Flatness is Necessary, Neural Collapse is Not: Rethinking Generalization via Grokking
von: Han, Ting, et al.
Veröffentlicht: (2025)
von: Han, Ting, et al.
Veröffentlicht: (2025)
Many-Objective Multi-Solution Transport
von: Li, Ziyue, et al.
Veröffentlicht: (2024)
von: Li, Ziyue, et al.
Veröffentlicht: (2024)
Topological Signatures of Grokking
von: Tang, Yifan, et al.
Veröffentlicht: (2026)
von: Tang, Yifan, et al.
Veröffentlicht: (2026)
The Pitfalls of Memorization: When Memorization Hurts Generalization
von: Bayat, Reza, et al.
Veröffentlicht: (2024)
von: Bayat, Reza, et al.
Veröffentlicht: (2024)
How Instruction and Reasoning Data shape Post-Training: Data Quality through the Lens of Layer-wise Gradients
von: Li, Ming, et al.
Veröffentlicht: (2025)
von: Li, Ming, et al.
Veröffentlicht: (2025)
V-REX: Benchmarking Exploratory Visual Reasoning via Chain-of-Questions
von: Fan, Chenrui, et al.
Veröffentlicht: (2025)
von: Fan, Chenrui, et al.
Veröffentlicht: (2025)
Variance-Adaptive Muon: Accelerating LLM Pretraining with NSR-Modulated and Variance-Scaled Momentum
von: Li, Jingru, et al.
Veröffentlicht: (2026)
von: Li, Jingru, et al.
Veröffentlicht: (2026)
Late-Stage Generalization Collapse in Grokking: Detecting anti-grokking with Weightwatcher
von: Prakash, Hari K, et al.
Veröffentlicht: (2026)
von: Prakash, Hari K, et al.
Veröffentlicht: (2026)
ILDR: Geometric Early Detection of Grokking
von: Golwala, Shreel
Veröffentlicht: (2026)
von: Golwala, Shreel
Veröffentlicht: (2026)
Grokking in Linear Estimators -- A Solvable Model that Groks without Understanding
von: Levi, Noam, et al.
Veröffentlicht: (2023)
von: Levi, Noam, et al.
Veröffentlicht: (2023)
Do Synthetic Trajectories Reflect Real Reward Hacking? A Systematic Study on Monitoring In-the-Wild Hacking in Code Generation
von: Li, Lichen, et al.
Veröffentlicht: (2026)
von: Li, Lichen, et al.
Veröffentlicht: (2026)
ColorBench: Can VLMs See and Understand the Colorful World? A Comprehensive Benchmark for Color Perception, Reasoning, and Robustness
von: Liang, Yijun, et al.
Veröffentlicht: (2025)
von: Liang, Yijun, et al.
Veröffentlicht: (2025)
Grokking and Generalization Collapse: Insights from \texttt{HTSR} theory
von: Prakash, Hari K., et al.
Veröffentlicht: (2025)
von: Prakash, Hari K., et al.
Veröffentlicht: (2025)
Membership and Memorization in LLM Knowledge Distillation
von: Zhang, Ziqi, et al.
Veröffentlicht: (2025)
von: Zhang, Ziqi, et al.
Veröffentlicht: (2025)
TNT: Improving Chunkwise Training for Test-Time Memorization
von: Li, Zeman, et al.
Veröffentlicht: (2025)
von: Li, Zeman, et al.
Veröffentlicht: (2025)
GrokAlign: Geometric Characterisation and Acceleration of Grokking
von: Walker, Thomas, et al.
Veröffentlicht: (2025)
von: Walker, Thomas, et al.
Veröffentlicht: (2025)
Breaking the Frozen Subspace: Importance Sampling for Low-Rank Optimization in LLM Pretraining
von: Zhang, Haochen, et al.
Veröffentlicht: (2025)
von: Zhang, Haochen, et al.
Veröffentlicht: (2025)
Memorizing Long-tail Data Can Help Generalization Through Composition
von: Zhou, Mo, et al.
Veröffentlicht: (2025)
von: Zhou, Mo, et al.
Veröffentlicht: (2025)
Pretrain-Test Task Alignment Governs Generalization in In-Context Learning
von: Letey, Mary I., et al.
Veröffentlicht: (2025)
von: Letey, Mary I., et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Skip a Layer or Loop it? Test-Time Depth Adaptation of Pretrained LLMs
von: Li, Ziyue, et al.
Veröffentlicht: (2025) -
Your Mixture-of-Experts LLM Is Secretly an Embedding Model For Free
von: Li, Ziyue, et al.
Veröffentlicht: (2024) -
R2-T2: Re-Routing in Test-Time for Multimodal Mixture-of-Experts
von: Li, Zhongyang, et al.
Veröffentlicht: (2025) -
Distributional Spectral Diagnostics for Localizing Grokking Transitions
von: Wang, Ziyue, et al.
Veröffentlicht: (2026) -
C3PO: Critical-Layer, Core-Expert, Collaborative Pathway Optimization for Test-Time Expert Re-Mixing
von: Li, Zhongyang, et al.
Veröffentlicht: (2025)