What Can Grokking Teach Us About Learning Under Nonstationarity?
Fuente:
arXiv
Saved in:
| Main Authors: | Lyle, Clare, Sokar, Gharda, Pascanu, Razvan, Gyorgy, Andras |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Deep Grokking: Would Deep Neural Networks Generalize Better?
by: Fan, Simin, et al.
Published: (2024)
by: Fan, Simin, et al.
Published: (2024)
Non-Stationary Learning of Neural Networks with Automatic Soft Parameter Reset
by: Galashov, Alexandre, et al.
Published: (2024)
by: Galashov, Alexandre, et al.
Published: (2024)
Disentangling the Causes of Plasticity Loss in Neural Networks
by: Lyle, Clare, et al.
Published: (2024)
by: Lyle, Clare, et al.
Published: (2024)
Lattice: Learning to Efficiently Compress the Memory
by: Karami, Mahdi, et al.
Published: (2025)
by: Karami, Mahdi, et al.
Published: (2025)
Fine-Tuned In-Context Learners for Efficient Adaptation
by: Bornschein, Jorg, et al.
Published: (2025)
by: Bornschein, Jorg, et al.
Published: (2025)
Normalization and effective learning rates in reinforcement learning
by: Lyle, Clare, et al.
Published: (2024)
by: Lyle, Clare, et al.
Published: (2024)
Revisiting Adam for Streaming Reinforcement Learning
by: Gogianu, Florin, et al.
Published: (2026)
by: Gogianu, Florin, et al.
Published: (2026)
Round and Round We Go! What makes Rotary Positional Encodings useful?
by: Barbero, Federico, et al.
Published: (2024)
by: Barbero, Federico, et al.
Published: (2024)
Revisiting Dynamic Evaluation: Online Adaptation for Large Language Models
by: Rannen-Triki, Amal, et al.
Published: (2024)
by: Rannen-Triki, Amal, et al.
Published: (2024)
NoProp: Training Neural Networks without Full Back-propagation or Full Forward-propagation
by: Li, Qinyu, et al.
Published: (2025)
by: Li, Qinyu, et al.
Published: (2025)
Latent Space Representations of Neural Algorithmic Reasoners
by: Mirjanić, Vladimir V., et al.
Published: (2023)
by: Mirjanić, Vladimir V., et al.
Published: (2023)
What Can 240,000 New Credit Transactions Tell Us About the Impact of NGEU Funds?
by: Ortiz, Alvaro, et al.
Published: (2025)
by: Ortiz, Alvaro, et al.
Published: (2025)
Meta-learning how to Share Credit among Macro-Actions
by: Hosu, Ionel-Alexandru, et al.
Published: (2025)
by: Hosu, Ionel-Alexandru, et al.
Published: (2025)
Mind the GAP! The Challenges of Scale in Pixel-based Deep Reinforcement Learning
by: Sokar, Ghada, et al.
Published: (2025)
by: Sokar, Ghada, et al.
Published: (2025)
Optimizers Qualitatively Alter Solutions And We Should Leverage This
by: Pascanu, Razvan, et al.
Published: (2025)
by: Pascanu, Razvan, et al.
Published: (2025)
Temporally Disentangled Representation Learning under Unknown Nonstationarity
by: Song, Xiangchen, et al.
Published: (2023)
by: Song, Xiangchen, et al.
Published: (2023)
Nonparametric Learning and Earning with One-Point Feedback under Nonstationarity
by: Yang, Xiangyu, et al.
Published: (2026)
by: Yang, Xiangyu, et al.
Published: (2026)
State Soup: In-Context Skill Learning, Retrieval and Mixing
by: Pióro, Maciej, et al.
Published: (2024)
by: Pióro, Maciej, et al.
Published: (2024)
MS-SSM: A Multi-Scale State Space Model for Efficient Sequence Modeling
by: Karami, Mahdi, et al.
Published: (2025)
by: Karami, Mahdi, et al.
Published: (2025)
No Representation, No Trust: Connecting Representation, Collapse, and Trust Issues in PPO
by: Moalla, Skander, et al.
Published: (2024)
by: Moalla, Skander, et al.
Published: (2024)
Attention as a Hypernetwork
by: Schug, Simon, et al.
Published: (2024)
by: Schug, Simon, et al.
Published: (2024)
Softmax is not Enough (for Sharp Size Generalisation)
by: Veličković, Petar, et al.
Published: (2024)
by: Veličković, Petar, et al.
Published: (2024)
Toward Understanding In-context vs. In-weight Learning
by: Chan, Bryan, et al.
Published: (2024)
by: Chan, Bryan, et al.
Published: (2024)
To Grok Grokking: Provable Grokking in Ridge Regression
by: Xu, Mingyue, et al.
Published: (2026)
by: Xu, Mingyue, et al.
Published: (2026)
The Nonstationarity-Complexity Tradeoff in Return Prediction
by: Capponi, Agostino, et al.
Published: (2025)
by: Capponi, Agostino, et al.
Published: (2025)
Hadamard product in deep learning: Introduction, Advances and Challenges
by: Chrysos, Grigorios G, et al.
Published: (2025)
by: Chrysos, Grigorios G, et al.
Published: (2025)
Beyond Statistical Learning: Exact Learning Is Essential for General Intelligence
by: György, András, et al.
Published: (2025)
by: György, András, et al.
Published: (2025)
Safe Continual Reinforcement Learning under Nonstationarity via Adaptive Safety Constraints
by: Tomashevskiy, Timofey
Published: (2026)
by: Tomashevskiy, Timofey
Published: (2026)
On the Peril of (Even a Little) Nonstationarity in Satisficing Regret Minimization
by: Zhang, Yixuan, et al.
Published: (2026)
by: Zhang, Yixuan, et al.
Published: (2026)
How Analysis Can Teach Us the Optimal Way to Design Neural Operators
by: Le, Vu-Anh, et al.
Published: (2024)
by: Le, Vu-Anh, et al.
Published: (2024)
LLMs are Greedy Agents: Effects of RL Fine-tuning on Decision-Making Abilities
by: Schmied, Thomas, et al.
Published: (2025)
by: Schmied, Thomas, et al.
Published: (2025)
The Complexity Dynamics of Grokking
by: DeMoss, Branton, et al.
Published: (2024)
by: DeMoss, Branton, et al.
Published: (2024)
Measuring Sharpness in Grokking
by: Miller, Jack, et al.
Published: (2024)
by: Miller, Jack, et al.
Published: (2024)
Fine-tuning Reinforcement Learning Models is Secretly a Forgetting Mitigation Problem
by: Wołczyk, Maciej, et al.
Published: (2024)
by: Wołczyk, Maciej, et al.
Published: (2024)
Perplexity Cannot Always Tell Right from Wrong
by: Veličković, Petar, et al.
Published: (2026)
by: Veličković, Petar, et al.
Published: (2026)
Bridging Lottery Ticket and Grokking: Understanding Grokking from Inner Structure of Networks
by: Minegishi, Gouki, et al.
Published: (2023)
by: Minegishi, Gouki, et al.
Published: (2023)
Unpacking Softmax: How Temperature Drives Representation Collapse, Compression, and Generalization
by: Masarczyk, Wojciech, et al.
Published: (2025)
by: Masarczyk, Wojciech, et al.
Published: (2025)
What Can Maslow Teach Us About Union Avoidance?
Published: (2025)
Published: (2025)
Universality of Linear Recurrences Followed by Non-linear Projections: Finite-Width Guarantees and Benefits of Complex Eigenvalues
by: Orvieto, Antonio, et al.
Published: (2023)
by: Orvieto, Antonio, et al.
Published: (2023)
TextTeacher: What Can Language Teach About Images?
by: Nauen, Tobias Christian, et al.
Published: (2026)
by: Nauen, Tobias Christian, et al.
Published: (2026)
Similar Items
-
Deep Grokking: Would Deep Neural Networks Generalize Better?
by: Fan, Simin, et al.
Published: (2024) -
Non-Stationary Learning of Neural Networks with Automatic Soft Parameter Reset
by: Galashov, Alexandre, et al.
Published: (2024) -
Disentangling the Causes of Plasticity Loss in Neural Networks
by: Lyle, Clare, et al.
Published: (2024) -
Lattice: Learning to Efficiently Compress the Memory
by: Karami, Mahdi, et al.
Published: (2025) -
Fine-Tuned In-Context Learners for Efficient Adaptation
by: Bornschein, Jorg, et al.
Published: (2025)