Grokking or Glitching? How Low-Precision Drives Slingshot Loss Spikes
Fuente:
arXiv
Salvato in:
| Autori principali: | Hanqing, Liu, Cao, Jianjun, Li, Yuanze, Zhou, Zijian |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Convex Dominance in Deep Learning I: A Scaling Law of Loss and Learning Rate
di: Bu, Zhiqi, et al.
Pubblicazione: (2026)
di: Bu, Zhiqi, et al.
Pubblicazione: (2026)
Adapprox: Adaptive Approximation in Adam Optimization via Randomized Low-Rank Matrices
di: Zhao, Pengxiang, et al.
Pubblicazione: (2024)
di: Zhao, Pengxiang, et al.
Pubblicazione: (2024)
When and How Unlabeled Data Provably Improve In-Context Learning
di: Li, Yingcong, et al.
Pubblicazione: (2025)
di: Li, Yingcong, et al.
Pubblicazione: (2025)
Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training
di: Liu, Hong, et al.
Pubblicazione: (2023)
di: Liu, Hong, et al.
Pubblicazione: (2023)
FOCUS: First Order Concentrated Updating Scheme
di: Liu, Yizhou, et al.
Pubblicazione: (2025)
di: Liu, Yizhou, et al.
Pubblicazione: (2025)
Wasserstein Distributionally Robust Regret Optimization for Reinforcement Learning from Human Feedback
di: Wang, Yikai, et al.
Pubblicazione: (2026)
di: Wang, Yikai, et al.
Pubblicazione: (2026)
On the Mechanism and Dynamics of Modular Addition: Fourier Features, Lottery Ticket, and Grokking
di: He, Jianliang, et al.
Pubblicazione: (2026)
di: He, Jianliang, et al.
Pubblicazione: (2026)
A Theoretical Framework for Grokking: Interpolation followed by Riemannian Norm Minimisation
di: Boursier, Etienne, et al.
Pubblicazione: (2025)
di: Boursier, Etienne, et al.
Pubblicazione: (2025)
On the Last-Iterate Convergence of Shuffling Gradient Methods
di: Liu, Zijian, et al.
Pubblicazione: (2024)
di: Liu, Zijian, et al.
Pubblicazione: (2024)
Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping
di: Liu, Zijian, et al.
Pubblicazione: (2024)
di: Liu, Zijian, et al.
Pubblicazione: (2024)
Improved Last-Iterate Convergence of Shuffling Gradient Methods for Nonsmooth Convex Optimization
di: Liu, Zijian, et al.
Pubblicazione: (2025)
di: Liu, Zijian, et al.
Pubblicazione: (2025)
Revisiting the Last-Iterate Convergence of Stochastic Gradient Methods
di: Liu, Zijian, et al.
Pubblicazione: (2023)
di: Liu, Zijian, et al.
Pubblicazione: (2023)
A Mechanism Study of Delayed Loss Spikes in Batch-Normalized Linear Models
di: Gao, Peifeng, et al.
Pubblicazione: (2026)
di: Gao, Peifeng, et al.
Pubblicazione: (2026)
Sign-Based Optimizers Are Effective Under Heavy-Tailed Noise
di: Yu, Dingzhi, et al.
Pubblicazione: (2026)
di: Yu, Dingzhi, et al.
Pubblicazione: (2026)
A Unified Understanding of Offline Data Selection and Online Self-refining Generation for Post-training LLMs
di: Xiao, Quan, et al.
Pubblicazione: (2025)
di: Xiao, Quan, et al.
Pubblicazione: (2025)
COS-DPO: Conditioned One-Shot Multi-Objective Fine-Tuning Framework
di: Ren, Yinuo, et al.
Pubblicazione: (2024)
di: Ren, Yinuo, et al.
Pubblicazione: (2024)
Distributional Surgery for Language Model Activations
di: Nguyen, Bao, et al.
Pubblicazione: (2025)
di: Nguyen, Bao, et al.
Pubblicazione: (2025)
Heavy-Tailed Class Imbalance and Why Adam Outperforms Gradient Descent on Language Models
di: Kunstner, Frederik, et al.
Pubblicazione: (2024)
di: Kunstner, Frederik, et al.
Pubblicazione: (2024)
SUMO: Subspace-Aware Moment-Orthogonalization for Accelerating Memory-Efficient LLM Training
di: Refael, Yehonathan, et al.
Pubblicazione: (2025)
di: Refael, Yehonathan, et al.
Pubblicazione: (2025)
Can Adaptive Gradient Methods Converge under Heavy-Tailed Noise? A Case Study of AdaGrad
di: Liu, Zijian
Pubblicazione: (2026)
di: Liu, Zijian
Pubblicazione: (2026)
In-Expectation Convergence of Stochastic Gradient Methods under Heavy-Tailed Noise
di: Liu, Zijian
Pubblicazione: (2026)
di: Liu, Zijian
Pubblicazione: (2026)
Clipped Gradient Methods for Nonsmooth Convex Optimization under Heavy-Tailed Noise: A Refined Analysis
di: Liu, Zijian
Pubblicazione: (2025)
di: Liu, Zijian
Pubblicazione: (2025)
Online Convex Optimization with Heavy Tails: Old Algorithms, New Regrets, and Applications
di: Liu, Zijian
Pubblicazione: (2025)
di: Liu, Zijian
Pubblicazione: (2025)
A Precise Characterization of SGD Stability Using Loss Surface Geometry
di: Dexter, Gregory, et al.
Pubblicazione: (2024)
di: Dexter, Gregory, et al.
Pubblicazione: (2024)
Variance-reduced Zeroth-Order Methods for Fine-Tuning Language Models
di: Gautam, Tanmay, et al.
Pubblicazione: (2024)
di: Gautam, Tanmay, et al.
Pubblicazione: (2024)
Understanding Forgetting in LLM Supervised Fine-Tuning and Preference Learning -- A Convex Optimization Perspective
di: Fernando, Heshan, et al.
Pubblicazione: (2024)
di: Fernando, Heshan, et al.
Pubblicazione: (2024)
GNMR: Runtime Stability Control for Low-Precision Large Language Model Training
di: Kong, Boao, et al.
Pubblicazione: (2026)
di: Kong, Boao, et al.
Pubblicazione: (2026)
Fine-grained Analysis of In-context Linear Estimation: Data, Architecture, and Beyond
di: Li, Yingcong, et al.
Pubblicazione: (2024)
di: Li, Yingcong, et al.
Pubblicazione: (2024)
Transformers as Support Vector Machines
di: Tarzanagh, Davoud Ataee, et al.
Pubblicazione: (2023)
di: Tarzanagh, Davoud Ataee, et al.
Pubblicazione: (2023)
LISA: Layerwise Importance Sampling for Memory-Efficient Large Language Model Fine-Tuning
di: Pan, Rui, et al.
Pubblicazione: (2024)
di: Pan, Rui, et al.
Pubblicazione: (2024)
One-Shot Safety Alignment for Large Language Models via Optimal Dualization
di: Huang, Xinmeng, et al.
Pubblicazione: (2024)
di: Huang, Xinmeng, et al.
Pubblicazione: (2024)
Gating is Weighting: Understanding Gated Linear Attention through In-context Learning
di: Li, Yingcong, et al.
Pubblicazione: (2025)
di: Li, Yingcong, et al.
Pubblicazione: (2025)
Mechanics of Next Token Prediction with Self-Attention
di: Li, Yingcong, et al.
Pubblicazione: (2024)
di: Li, Yingcong, et al.
Pubblicazione: (2024)
Solving General Natural-Language-Description Optimization Problems with Large Language Models
di: Zhang, Jihai, et al.
Pubblicazione: (2024)
di: Zhang, Jihai, et al.
Pubblicazione: (2024)
StablePCA: Distributionally Robust Learning of Shared Representations from Multi-Source Data
di: Wang, Zhenyu, et al.
Pubblicazione: (2025)
di: Wang, Zhenyu, et al.
Pubblicazione: (2025)
DiaBlo: Diagonal Blocks Are Sufficient For Finetuning
di: Gurses, Selcuk, et al.
Pubblicazione: (2025)
di: Gurses, Selcuk, et al.
Pubblicazione: (2025)
OptiChat: Bridging Optimization Models and Practitioners with Large Language Models
di: Chen, Hao, et al.
Pubblicazione: (2025)
di: Chen, Hao, et al.
Pubblicazione: (2025)
BOOOM: Loss-Function-Agnostic Black-Box Optimization over Orthonormal Manifolds for Machine Learning and Statistical Inference
di: Kim, Beomchang, et al.
Pubblicazione: (2026)
di: Kim, Beomchang, et al.
Pubblicazione: (2026)
How Multimodal Integration Boost the Performance of LLM for Optimization: Case Study on Capacitated Vehicle Routing Problems
di: Huang, Yuxiao, et al.
Pubblicazione: (2024)
di: Huang, Yuxiao, et al.
Pubblicazione: (2024)
When and Why SignSGD Outperforms SGD: A Theoretical Study Based on $\ell_1$-norm Lower Bounds
di: Tao, Hongyi, et al.
Pubblicazione: (2026)
di: Tao, Hongyi, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Convex Dominance in Deep Learning I: A Scaling Law of Loss and Learning Rate
di: Bu, Zhiqi, et al.
Pubblicazione: (2026) -
Adapprox: Adaptive Approximation in Adam Optimization via Randomized Low-Rank Matrices
di: Zhao, Pengxiang, et al.
Pubblicazione: (2024) -
When and How Unlabeled Data Provably Improve In-Context Learning
di: Li, Yingcong, et al.
Pubblicazione: (2025) -
Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training
di: Liu, Hong, et al.
Pubblicazione: (2023) -
FOCUS: First Order Concentrated Updating Scheme
di: Liu, Yizhou, et al.
Pubblicazione: (2025)