Grokking or Glitching? How Low-Precision Drives Slingshot Loss Spikes
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hanqing, Liu, Cao, Jianjun, Li, Yuanze, Zhou, Zijian |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Convex Dominance in Deep Learning I: A Scaling Law of Loss and Learning Rate
von: Bu, Zhiqi, et al.
Veröffentlicht: (2026)
von: Bu, Zhiqi, et al.
Veröffentlicht: (2026)
Adapprox: Adaptive Approximation in Adam Optimization via Randomized Low-Rank Matrices
von: Zhao, Pengxiang, et al.
Veröffentlicht: (2024)
von: Zhao, Pengxiang, et al.
Veröffentlicht: (2024)
When and How Unlabeled Data Provably Improve In-Context Learning
von: Li, Yingcong, et al.
Veröffentlicht: (2025)
von: Li, Yingcong, et al.
Veröffentlicht: (2025)
Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training
von: Liu, Hong, et al.
Veröffentlicht: (2023)
von: Liu, Hong, et al.
Veröffentlicht: (2023)
FOCUS: First Order Concentrated Updating Scheme
von: Liu, Yizhou, et al.
Veröffentlicht: (2025)
von: Liu, Yizhou, et al.
Veröffentlicht: (2025)
Wasserstein Distributionally Robust Regret Optimization for Reinforcement Learning from Human Feedback
von: Wang, Yikai, et al.
Veröffentlicht: (2026)
von: Wang, Yikai, et al.
Veröffentlicht: (2026)
On the Mechanism and Dynamics of Modular Addition: Fourier Features, Lottery Ticket, and Grokking
von: He, Jianliang, et al.
Veröffentlicht: (2026)
von: He, Jianliang, et al.
Veröffentlicht: (2026)
A Theoretical Framework for Grokking: Interpolation followed by Riemannian Norm Minimisation
von: Boursier, Etienne, et al.
Veröffentlicht: (2025)
von: Boursier, Etienne, et al.
Veröffentlicht: (2025)
On the Last-Iterate Convergence of Shuffling Gradient Methods
von: Liu, Zijian, et al.
Veröffentlicht: (2024)
von: Liu, Zijian, et al.
Veröffentlicht: (2024)
Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping
von: Liu, Zijian, et al.
Veröffentlicht: (2024)
von: Liu, Zijian, et al.
Veröffentlicht: (2024)
Improved Last-Iterate Convergence of Shuffling Gradient Methods for Nonsmooth Convex Optimization
von: Liu, Zijian, et al.
Veröffentlicht: (2025)
von: Liu, Zijian, et al.
Veröffentlicht: (2025)
Revisiting the Last-Iterate Convergence of Stochastic Gradient Methods
von: Liu, Zijian, et al.
Veröffentlicht: (2023)
von: Liu, Zijian, et al.
Veröffentlicht: (2023)
A Mechanism Study of Delayed Loss Spikes in Batch-Normalized Linear Models
von: Gao, Peifeng, et al.
Veröffentlicht: (2026)
von: Gao, Peifeng, et al.
Veröffentlicht: (2026)
Sign-Based Optimizers Are Effective Under Heavy-Tailed Noise
von: Yu, Dingzhi, et al.
Veröffentlicht: (2026)
von: Yu, Dingzhi, et al.
Veröffentlicht: (2026)
A Unified Understanding of Offline Data Selection and Online Self-refining Generation for Post-training LLMs
von: Xiao, Quan, et al.
Veröffentlicht: (2025)
von: Xiao, Quan, et al.
Veröffentlicht: (2025)
COS-DPO: Conditioned One-Shot Multi-Objective Fine-Tuning Framework
von: Ren, Yinuo, et al.
Veröffentlicht: (2024)
von: Ren, Yinuo, et al.
Veröffentlicht: (2024)
Distributional Surgery for Language Model Activations
von: Nguyen, Bao, et al.
Veröffentlicht: (2025)
von: Nguyen, Bao, et al.
Veröffentlicht: (2025)
Heavy-Tailed Class Imbalance and Why Adam Outperforms Gradient Descent on Language Models
von: Kunstner, Frederik, et al.
Veröffentlicht: (2024)
von: Kunstner, Frederik, et al.
Veröffentlicht: (2024)
SUMO: Subspace-Aware Moment-Orthogonalization for Accelerating Memory-Efficient LLM Training
von: Refael, Yehonathan, et al.
Veröffentlicht: (2025)
von: Refael, Yehonathan, et al.
Veröffentlicht: (2025)
Can Adaptive Gradient Methods Converge under Heavy-Tailed Noise? A Case Study of AdaGrad
von: Liu, Zijian
Veröffentlicht: (2026)
von: Liu, Zijian
Veröffentlicht: (2026)
In-Expectation Convergence of Stochastic Gradient Methods under Heavy-Tailed Noise
von: Liu, Zijian
Veröffentlicht: (2026)
von: Liu, Zijian
Veröffentlicht: (2026)
Clipped Gradient Methods for Nonsmooth Convex Optimization under Heavy-Tailed Noise: A Refined Analysis
von: Liu, Zijian
Veröffentlicht: (2025)
von: Liu, Zijian
Veröffentlicht: (2025)
Online Convex Optimization with Heavy Tails: Old Algorithms, New Regrets, and Applications
von: Liu, Zijian
Veröffentlicht: (2025)
von: Liu, Zijian
Veröffentlicht: (2025)
A Precise Characterization of SGD Stability Using Loss Surface Geometry
von: Dexter, Gregory, et al.
Veröffentlicht: (2024)
von: Dexter, Gregory, et al.
Veröffentlicht: (2024)
Variance-reduced Zeroth-Order Methods for Fine-Tuning Language Models
von: Gautam, Tanmay, et al.
Veröffentlicht: (2024)
von: Gautam, Tanmay, et al.
Veröffentlicht: (2024)
Understanding Forgetting in LLM Supervised Fine-Tuning and Preference Learning -- A Convex Optimization Perspective
von: Fernando, Heshan, et al.
Veröffentlicht: (2024)
von: Fernando, Heshan, et al.
Veröffentlicht: (2024)
GNMR: Runtime Stability Control for Low-Precision Large Language Model Training
von: Kong, Boao, et al.
Veröffentlicht: (2026)
von: Kong, Boao, et al.
Veröffentlicht: (2026)
Fine-grained Analysis of In-context Linear Estimation: Data, Architecture, and Beyond
von: Li, Yingcong, et al.
Veröffentlicht: (2024)
von: Li, Yingcong, et al.
Veröffentlicht: (2024)
Transformers as Support Vector Machines
von: Tarzanagh, Davoud Ataee, et al.
Veröffentlicht: (2023)
von: Tarzanagh, Davoud Ataee, et al.
Veröffentlicht: (2023)
LISA: Layerwise Importance Sampling for Memory-Efficient Large Language Model Fine-Tuning
von: Pan, Rui, et al.
Veröffentlicht: (2024)
von: Pan, Rui, et al.
Veröffentlicht: (2024)
One-Shot Safety Alignment for Large Language Models via Optimal Dualization
von: Huang, Xinmeng, et al.
Veröffentlicht: (2024)
von: Huang, Xinmeng, et al.
Veröffentlicht: (2024)
Gating is Weighting: Understanding Gated Linear Attention through In-context Learning
von: Li, Yingcong, et al.
Veröffentlicht: (2025)
von: Li, Yingcong, et al.
Veröffentlicht: (2025)
Mechanics of Next Token Prediction with Self-Attention
von: Li, Yingcong, et al.
Veröffentlicht: (2024)
von: Li, Yingcong, et al.
Veröffentlicht: (2024)
Solving General Natural-Language-Description Optimization Problems with Large Language Models
von: Zhang, Jihai, et al.
Veröffentlicht: (2024)
von: Zhang, Jihai, et al.
Veröffentlicht: (2024)
StablePCA: Distributionally Robust Learning of Shared Representations from Multi-Source Data
von: Wang, Zhenyu, et al.
Veröffentlicht: (2025)
von: Wang, Zhenyu, et al.
Veröffentlicht: (2025)
DiaBlo: Diagonal Blocks Are Sufficient For Finetuning
von: Gurses, Selcuk, et al.
Veröffentlicht: (2025)
von: Gurses, Selcuk, et al.
Veröffentlicht: (2025)
OptiChat: Bridging Optimization Models and Practitioners with Large Language Models
von: Chen, Hao, et al.
Veröffentlicht: (2025)
von: Chen, Hao, et al.
Veröffentlicht: (2025)
BOOOM: Loss-Function-Agnostic Black-Box Optimization over Orthonormal Manifolds for Machine Learning and Statistical Inference
von: Kim, Beomchang, et al.
Veröffentlicht: (2026)
von: Kim, Beomchang, et al.
Veröffentlicht: (2026)
How Multimodal Integration Boost the Performance of LLM for Optimization: Case Study on Capacitated Vehicle Routing Problems
von: Huang, Yuxiao, et al.
Veröffentlicht: (2024)
von: Huang, Yuxiao, et al.
Veröffentlicht: (2024)
When and Why SignSGD Outperforms SGD: A Theoretical Study Based on $\ell_1$-norm Lower Bounds
von: Tao, Hongyi, et al.
Veröffentlicht: (2026)
von: Tao, Hongyi, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Convex Dominance in Deep Learning I: A Scaling Law of Loss and Learning Rate
von: Bu, Zhiqi, et al.
Veröffentlicht: (2026) -
Adapprox: Adaptive Approximation in Adam Optimization via Randomized Low-Rank Matrices
von: Zhao, Pengxiang, et al.
Veröffentlicht: (2024) -
When and How Unlabeled Data Provably Improve In-Context Learning
von: Li, Yingcong, et al.
Veröffentlicht: (2025) -
Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training
von: Liu, Hong, et al.
Veröffentlicht: (2023) -
FOCUS: First Order Concentrated Updating Scheme
von: Liu, Yizhou, et al.
Veröffentlicht: (2025)