Adaptive Preconditioners Trigger Loss Spikes in Adam
Fuente:
arXiv
Saved in:
| Main Authors: | Bai, Zhiwei, Zhou, Zhangchen, Zhao, Jiajie, Li, Xiaolong, Li, Zhiyu, Xiong, Feiyu, Yang, Hongkang, Zhang, Yaoyu, Xu, Zhi-Qin John |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
An overview of condensation phenomenon in deep learning
by: Xu, Zhi-Qin John, et al.
Published: (2025)
by: Xu, Zhi-Qin John, et al.
Published: (2025)
A rationale from frequency perspective for grokking in training neural network
by: Zhou, Zhangchen, et al.
Published: (2024)
by: Zhou, Zhangchen, et al.
Published: (2024)
Loss Spike in Training Neural Networks
by: Li, Xiaolong, et al.
Published: (2023)
by: Li, Xiaolong, et al.
Published: (2023)
Embedding Principle in Depth for the Loss Landscape Analysis of Deep Neural Networks
by: Bai, Zhiwei, et al.
Published: (2022)
by: Bai, Zhiwei, et al.
Published: (2022)
Connectivity Shapes Implicit Regularization in Matrix Factorization Models for Matrix Completion
by: Bai, Zhiwei, et al.
Published: (2024)
by: Bai, Zhiwei, et al.
Published: (2024)
Disentangle Sample Size and Initialization Effect on Perfect Generalization for Single-Neuron Target
by: Zhao, Jiajie, et al.
Published: (2024)
by: Zhao, Jiajie, et al.
Published: (2024)
Scalable Complexity Control Facilitates Reasoning Ability of LLMs
by: Hang, Liangkai, et al.
Published: (2025)
by: Hang, Liangkai, et al.
Published: (2025)
Anchor function: a type of benchmark functions for studying language models
by: Zhang, Zhongwang, et al.
Published: (2024)
by: Zhang, Zhongwang, et al.
Published: (2024)
Understanding the Language Model to Solve the Symbolic Multi-Step Reasoning Problem from the Perspective of Buffer Mechanism
by: Wang, Zhiwei, et al.
Published: (2024)
by: Wang, Zhiwei, et al.
Published: (2024)
Initialization is Critical to Whether Transformers Fit Composite Functions by Reasoning or Memorizing
by: Zhang, Zhongwang, et al.
Published: (2024)
by: Zhang, Zhongwang, et al.
Published: (2024)
Loss Jump During Loss Switch in Solving PDEs with Neural Networks
by: Wang, Zhiwei, et al.
Published: (2024)
by: Wang, Zhiwei, et al.
Published: (2024)
Complexity Control Facilitates Reasoning-Based Compositional Generalization in Transformers
by: Zhang, Zhongwang, et al.
Published: (2025)
by: Zhang, Zhongwang, et al.
Published: (2025)
Overview frequency principle/spectral bias in deep learning
by: Xu, Zhi-Qin John, et al.
Published: (2022)
by: Xu, Zhi-Qin John, et al.
Published: (2022)
Phase Diagram of Initial Condensation for Two-layer Neural Networks
by: Chen, Zhengan, et al.
Published: (2023)
by: Chen, Zhengan, et al.
Published: (2023)
SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam
by: Peng, Hanyang, et al.
Published: (2025)
by: Peng, Hanyang, et al.
Published: (2025)
Structured Preconditioners in Adaptive Optimization: A Unified Analysis
by: Xie, Shuo, et al.
Published: (2025)
by: Xie, Shuo, et al.
Published: (2025)
Local Linear Recovery Guarantee of Deep Neural Networks at Overparameterization
by: Zhang, Yaoyu, et al.
Published: (2024)
by: Zhang, Yaoyu, et al.
Published: (2024)
Dynamic Regret via Discounted-to-Dynamic Reduction with Applications to Curved Losses and Adam Optimizer
by: Xie, Yan-Feng, et al.
Published: (2026)
by: Xie, Yan-Feng, et al.
Published: (2026)
Adam Exploits $\ell_\infty$-geometry of Loss Landscape via Coordinate-wise Adaptivity
by: Xie, Shuo, et al.
Published: (2024)
by: Xie, Shuo, et al.
Published: (2024)
Improving Adaptive Moment Optimization via Preconditioner Diagonalization
by: Nguyen, Son, et al.
Published: (2025)
by: Nguyen, Son, et al.
Published: (2025)
xMLP: Revolutionizing Private Inference with Exclusive Square Activation
by: Li, Jiajie, et al.
Published: (2024)
by: Li, Jiajie, et al.
Published: (2024)
Understanding Adam Optimizer via Online Learning of Updates: Adam is FTRL in Disguise
by: Ahn, Kwangjun, et al.
Published: (2024)
by: Ahn, Kwangjun, et al.
Published: (2024)
Neural Network Based Framework for Passive Intermodulation Cancellation in MIMO Systems
by: Li, Xiaolong, et al.
Published: (2025)
by: Li, Xiaolong, et al.
Published: (2025)
Provable Adaptivity of Adam under Non-uniform Smoothness
by: Wang, Bohan, et al.
Published: (2022)
by: Wang, Bohan, et al.
Published: (2022)
Grokking or Glitching? How Low-Precision Drives Slingshot Loss Spikes
by: Hanqing, Liu, et al.
Published: (2026)
by: Hanqing, Liu, et al.
Published: (2026)
Learning Sparse Approximate Inverse Preconditioners for Conjugate Gradient Solvers on GPUs
by: Yang, Zherui, et al.
Published: (2025)
by: Yang, Zherui, et al.
Published: (2025)
UT-ACA: Uncertainty-Triggered Adaptive Context Allocation for Long-Context Inference
by: Zhou, Lang, et al.
Published: (2026)
by: Zhou, Lang, et al.
Published: (2026)
Generative modeling of Sparse Approximate Inverse Preconditioners
by: Li, Mou, et al.
Published: (2024)
by: Li, Mou, et al.
Published: (2024)
A Computational Approach to Improving Fairness in K-means Clustering
by: Zhou, Guancheng, et al.
Published: (2025)
by: Zhou, Guancheng, et al.
Published: (2025)
Frequency Principle: Fourier Analysis Sheds Light on Deep Neural Networks
by: Xu, Zhi-Qin John, et al.
Published: (2019)
by: Xu, Zhi-Qin John, et al.
Published: (2019)
Achilles' Heel of Mamba: Essential difficulties of the Mamba architecture demonstrated by synthetic data
by: Chen, Tianyi, et al.
Published: (2025)
by: Chen, Tianyi, et al.
Published: (2025)
KodCode: A Diverse, Challenging, and Verifiable Synthetic Dataset for Coding
by: Xu, Zhangchen, et al.
Published: (2025)
by: Xu, Zhangchen, et al.
Published: (2025)
Theoretical Learning Performance of Graph Neural Networks: The Impact of Jumping Connections and Layer-wise Sparsification
by: Sun, Jiawei, et al.
Published: (2025)
by: Sun, Jiawei, et al.
Published: (2025)
Transformers Learn the Optimal DDPM Denoiser for Multi-Token GMMs
by: Li, Hongkang, et al.
Published: (2026)
by: Li, Hongkang, et al.
Published: (2026)
Curvature-Informed SGD via General Purpose Lie-Group Preconditioners
by: Pooladzandi, Omead, et al.
Published: (2024)
by: Pooladzandi, Omead, et al.
Published: (2024)
Anon: Extrapolating Adaptivity Beyond SGD and Adam
by: Zhang, Yiheng, et al.
Published: (2026)
by: Zhang, Yiheng, et al.
Published: (2026)
Towards Reliable Pediatric Brain Tumor Segmentation: Task-Specific nnU-Net Enhancements
by: Li, Xiaolong, et al.
Published: (2025)
by: Li, Xiaolong, et al.
Published: (2025)
Spectrally-Corrected and Regularized Linear Discriminant Analysis for Spiked Covariance Model
by: Li, Hua, et al.
Published: (2022)
by: Li, Hua, et al.
Published: (2022)
SPAM: Spike-Aware Adam with Momentum Reset for Stable LLM Training
by: Huang, Tianjin, et al.
Published: (2025)
by: Huang, Tianjin, et al.
Published: (2025)
Generalization Error of GAN from the Discriminator's Perspective
by: Yang, Hongkang, et al.
Published: (2021)
by: Yang, Hongkang, et al.
Published: (2021)
Similar Items
-
An overview of condensation phenomenon in deep learning
by: Xu, Zhi-Qin John, et al.
Published: (2025) -
A rationale from frequency perspective for grokking in training neural network
by: Zhou, Zhangchen, et al.
Published: (2024) -
Loss Spike in Training Neural Networks
by: Li, Xiaolong, et al.
Published: (2023) -
Embedding Principle in Depth for the Loss Landscape Analysis of Deep Neural Networks
by: Bai, Zhiwei, et al.
Published: (2022) -
Connectivity Shapes Implicit Regularization in Matrix Factorization Models for Matrix Completion
by: Bai, Zhiwei, et al.
Published: (2024)