Loss Spike in Training Neural Networks
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Xiaolong, Xu, Zhi-Qin John, Zhang, Zhongwang |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Loss Jump During Loss Switch in Solving PDEs with Neural Networks
by: Wang, Zhiwei, et al.
Published: (2024)
by: Wang, Zhiwei, et al.
Published: (2024)
Reasoning Bias of Next Token Prediction Training
by: Lin, Pengxiao, et al.
Published: (2025)
by: Lin, Pengxiao, et al.
Published: (2025)
Adaptive Preconditioners Trigger Loss Spikes in Adam
by: Bai, Zhiwei, et al.
Published: (2025)
by: Bai, Zhiwei, et al.
Published: (2025)
An Analysis for Reasoning Bias of Language Models with Small Initialization
by: Yao, Junjie, et al.
Published: (2025)
by: Yao, Junjie, et al.
Published: (2025)
Embedding Principle in Depth for the Loss Landscape Analysis of Deep Neural Networks
by: Bai, Zhiwei, et al.
Published: (2022)
by: Bai, Zhiwei, et al.
Published: (2022)
Anchor function: a type of benchmark functions for studying language models
by: Zhang, Zhongwang, et al.
Published: (2024)
by: Zhang, Zhongwang, et al.
Published: (2024)
Initialization is Critical to Whether Transformers Fit Composite Functions by Reasoning or Memorizing
by: Zhang, Zhongwang, et al.
Published: (2024)
by: Zhang, Zhongwang, et al.
Published: (2024)
Local Linear Recovery Guarantee of Deep Neural Networks at Overparameterization
by: Zhang, Yaoyu, et al.
Published: (2024)
by: Zhang, Yaoyu, et al.
Published: (2024)
Neural Network Based Framework for Passive Intermodulation Cancellation in MIMO Systems
by: Li, Xiaolong, et al.
Published: (2025)
by: Li, Xiaolong, et al.
Published: (2025)
Complexity Control Facilitates Reasoning-Based Compositional Generalization in Transformers
by: Zhang, Zhongwang, et al.
Published: (2025)
by: Zhang, Zhongwang, et al.
Published: (2025)
A Closer Look at Knowledge Distillation in Spiking Neural Network Training
by: Liu, Xu, et al.
Published: (2025)
by: Liu, Xu, et al.
Published: (2025)
Training of Spiking Neural Networks with Expectation-Propagation
by: Yao, Dan, et al.
Published: (2025)
by: Yao, Dan, et al.
Published: (2025)
Efficient and Flexible Method for Reducing Moderate-size Deep Neural Networks with Condensation
by: Chen, Tianyi, et al.
Published: (2024)
by: Chen, Tianyi, et al.
Published: (2024)
Spiking Brain Compression: Post-Training Second-order Compression for Spiking Neural Networks
by: Shi, Lianfeng, et al.
Published: (2025)
by: Shi, Lianfeng, et al.
Published: (2025)
A Principled Bayesian Framework for Training Binary and Spiking Neural Networks
by: Walker, James A., et al.
Published: (2025)
by: Walker, James A., et al.
Published: (2025)
Sharpness Aware Surrogate Training for Spiking Neural Networks
by: Nicholson, Maximilian
Published: (2026)
by: Nicholson, Maximilian
Published: (2026)
Bullet Trains: Parallelizing Training of Temporally Precise Spiking Neural Networks
by: Morrill, Todd, et al.
Published: (2026)
by: Morrill, Todd, et al.
Published: (2026)
Z-Error Loss for Training Neural Networks
by: Godin, Guillaume
Published: (2025)
by: Godin, Guillaume
Published: (2025)
SpikeVoice: High-Quality Text-to-Speech Via Efficient Spiking Neural Network
by: Wang, Kexin, et al.
Published: (2024)
by: Wang, Kexin, et al.
Published: (2024)
Spectral Alignment as Predictor of Loss Explosion in Neural Network Training
by: Qiu, Haiquan, et al.
Published: (2025)
by: Qiu, Haiquan, et al.
Published: (2025)
Exploring the Potential of Spiking Neural Networks in UWB Channel Estimation
by: Zhang, Youdong, et al.
Published: (2025)
by: Zhang, Youdong, et al.
Published: (2025)
Spiking Graph Neural Network on Riemannian Manifolds
by: Sun, Li, et al.
Published: (2024)
by: Sun, Li, et al.
Published: (2024)
Training a General Spiking Neural Network with Improved Efficiency and Minimum Latency
by: Yao, Yunpeng, et al.
Published: (2024)
by: Yao, Yunpeng, et al.
Published: (2024)
Energy-Efficient and Dequantization-Free Q-LLMs: A Spiking Neural Network Approach to Salient Value Mitigation
by: Wang, Chenyu, et al.
Published: (2025)
by: Wang, Chenyu, et al.
Published: (2025)
ChronoSpike: An Adaptive Spiking Graph Neural Network for Dynamic Graphs
by: Jahin, Md Abrar, et al.
Published: (2026)
by: Jahin, Md Abrar, et al.
Published: (2026)
Online Pseudo-Zeroth-Order Training of Neuromorphic Spiking Neural Networks
by: Xiao, Mingqing, et al.
Published: (2024)
by: Xiao, Mingqing, et al.
Published: (2024)
Training Spiking Neural Networks via Augmented Direct Feedback Alignment
by: Zhang, Yongbo, et al.
Published: (2024)
by: Zhang, Yongbo, et al.
Published: (2024)
SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam
by: Peng, Hanyang, et al.
Published: (2025)
by: Peng, Hanyang, et al.
Published: (2025)
SQUAT: Stateful Quantization-Aware Training in Recurrent Spiking Neural Networks
by: Venkatesh, Sreyes, et al.
Published: (2024)
by: Venkatesh, Sreyes, et al.
Published: (2024)
Visualizing, Rethinking, and Mining the Loss Landscape of Deep Neural Networks
by: Xu, Yichu, et al.
Published: (2024)
by: Xu, Yichu, et al.
Published: (2024)
Advancing Training Efficiency of Deep Spiking Neural Networks through Rate-based Backpropagation
by: Yu, Chengting, et al.
Published: (2024)
by: Yu, Chengting, et al.
Published: (2024)
On the Adversarial Robustness of Spiking Neural Networks Trained by Local Learning
by: Lin, Jiaqi, et al.
Published: (2025)
by: Lin, Jiaqi, et al.
Published: (2025)
ADMM-Based Training for Spiking Neural Networks
by: Perin, Giovanni, et al.
Published: (2025)
by: Perin, Giovanni, et al.
Published: (2025)
SpikingGamma: Surrogate-Gradient Free and Temporally Precise Online Training of Spiking Neural Networks with Smoothed Delays
by: Koopman, Roel, et al.
Published: (2026)
by: Koopman, Roel, et al.
Published: (2026)
Quantization of Spiking Neural Networks Beyond Accuracy
by: Smith, Evan Gibson, et al.
Published: (2026)
by: Smith, Evan Gibson, et al.
Published: (2026)
Congestion-Aware Dynamic Axonal Delay for Spiking Neural Networks
by: Bai, Dewei, et al.
Published: (2026)
by: Bai, Dewei, et al.
Published: (2026)
Physics-Informed Spiking Neural Networks via Conservative Flux Quantization
by: Zhang, Chi, et al.
Published: (2025)
by: Zhang, Chi, et al.
Published: (2025)
Neuronal Self-Adaptation Enhances Capacity and Robustness of Representation in Spiking Neural Networks
by: Yang, Zhuobin, et al.
Published: (2026)
by: Yang, Zhuobin, et al.
Published: (2026)
SpikeLLM: Scaling up Spiking Neural Network to Large Language Models via Saliency-based Spiking
by: Xing, Xingrun, et al.
Published: (2024)
by: Xing, Xingrun, et al.
Published: (2024)
A Quasi-Wasserstein Loss for Learning Graph Neural Networks
by: Cheng, Minjie, et al.
Published: (2023)
by: Cheng, Minjie, et al.
Published: (2023)
Similar Items
-
Loss Jump During Loss Switch in Solving PDEs with Neural Networks
by: Wang, Zhiwei, et al.
Published: (2024) -
Reasoning Bias of Next Token Prediction Training
by: Lin, Pengxiao, et al.
Published: (2025) -
Adaptive Preconditioners Trigger Loss Spikes in Adam
by: Bai, Zhiwei, et al.
Published: (2025) -
An Analysis for Reasoning Bias of Language Models with Small Initialization
by: Yao, Junjie, et al.
Published: (2025) -
Embedding Principle in Depth for the Loss Landscape Analysis of Deep Neural Networks
by: Bai, Zhiwei, et al.
Published: (2022)