Why Gradients Rapidly Increase Near the End of Training
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Defazio, Aaron |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ScheduleFree+: Scaling Learning-Rate-Free & Schedule-Free Learning to Large Language Models
von: Defazio, Aaron
Veröffentlicht: (2026)
von: Defazio, Aaron
Veröffentlicht: (2026)
Prodigy: An Expeditiously Adaptive Parameter-Free Learner
von: Mishchenko, Konstantin, et al.
Veröffentlicht: (2023)
von: Mishchenko, Konstantin, et al.
Veröffentlicht: (2023)
Smoothing DiLoCo with Primal Averaging for Faster Training of LLMs
von: Defazio, Aaron, et al.
Veröffentlicht: (2025)
von: Defazio, Aaron, et al.
Veröffentlicht: (2025)
Optimal Linear Decay Learning Rate Schedules and Further Refinements
von: Defazio, Aaron, et al.
Veröffentlicht: (2023)
von: Defazio, Aaron, et al.
Veröffentlicht: (2023)
Purifying Shampoo: Investigating Shampoo's Heuristics by Decomposing its Preconditioner
von: Eschenhagen, Runa, et al.
Veröffentlicht: (2025)
von: Eschenhagen, Runa, et al.
Veröffentlicht: (2025)
Conditional Rectified Flow-based End-to-End Rapid Seismic Inversion Method
von: Xu, Haofei, et al.
Veröffentlicht: (2026)
von: Xu, Haofei, et al.
Veröffentlicht: (2026)
The Road Less Scheduled
von: Defazio, Aaron, et al.
Veröffentlicht: (2024)
von: Defazio, Aaron, et al.
Veröffentlicht: (2024)
End-to-End On-Device Quantization-Aware Training for LLMs at Inference Cost
von: Tan, Qitao, et al.
Veröffentlicht: (2025)
von: Tan, Qitao, et al.
Veröffentlicht: (2025)
Learning to Remember: End-to-End Training of Memory Agents for Long-Context Reasoning
von: Zhang, Kehao, et al.
Veröffentlicht: (2026)
von: Zhang, Kehao, et al.
Veröffentlicht: (2026)
Why Does Stochastic Gradient Descent Slow Down in Low-Precision Training?
von: Yun, Vincent-Daniel
Veröffentlicht: (2025)
von: Yun, Vincent-Daniel
Veröffentlicht: (2025)
Revisiting the Relationship between Adversarial and Clean Training: Why Clean Training Can Make Adversarial Training Better
von: Zhou, MingWei, et al.
Veröffentlicht: (2025)
von: Zhou, MingWei, et al.
Veröffentlicht: (2025)
L-MoE: End-to-End Training of a Lightweight Mixture of Low-Rank Adaptation Experts
von: Ji, Shihao, et al.
Veröffentlicht: (2025)
von: Ji, Shihao, et al.
Veröffentlicht: (2025)
Why Inference in Large Models Becomes Decomposable After Training
von: Jin, Jidong
Veröffentlicht: (2026)
von: Jin, Jidong
Veröffentlicht: (2026)
Gradient-Congruity Guided Federated Sparse Training
von: Tian, Chris Xing, et al.
Veröffentlicht: (2024)
von: Tian, Chris Xing, et al.
Veröffentlicht: (2024)
Gradient-Free Training of Quantized Neural Networks
von: Cohen, Noa, et al.
Veröffentlicht: (2024)
von: Cohen, Noa, et al.
Veröffentlicht: (2024)
Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention
von: Qiu, Haiquan, et al.
Veröffentlicht: (2025)
von: Qiu, Haiquan, et al.
Veröffentlicht: (2025)
Gradient Multi-Normalization for Stateless and Scalable LLM Training
von: Scetbon, Meyer, et al.
Veröffentlicht: (2025)
von: Scetbon, Meyer, et al.
Veröffentlicht: (2025)
Certification for Differentially Private Prediction in Gradient-Based Training
von: Wicker, Matthew, et al.
Veröffentlicht: (2024)
von: Wicker, Matthew, et al.
Veröffentlicht: (2024)
GIO: Gradient Information Optimization for Training Dataset Selection
von: Everaert, Dante, et al.
Veröffentlicht: (2023)
von: Everaert, Dante, et al.
Veröffentlicht: (2023)
Why Adam Works Better with $β_1 = β_2$: The Missing Gradient Scale Invariance Principle
von: Fernández-Hernández, Alberto, et al.
Veröffentlicht: (2026)
von: Fernández-Hernández, Alberto, et al.
Veröffentlicht: (2026)
Gradient Inversion Transcript: Leveraging Robust Generative Priors to Reconstruct Training Data from Gradient Leakage
von: Chen, Xinping, et al.
Veröffentlicht: (2025)
von: Chen, Xinping, et al.
Veröffentlicht: (2025)
Predictive Concept Decoders: Training Scalable End-to-End Interpretability Assistants
von: Huang, Vincent, et al.
Veröffentlicht: (2025)
von: Huang, Vincent, et al.
Veröffentlicht: (2025)
Uncovering Gradient Inversion Risks in Practical Language Model Training
von: Feng, Xinguo, et al.
Veröffentlicht: (2025)
von: Feng, Xinguo, et al.
Veröffentlicht: (2025)
Imbalanced Gradients in RL Post-Training of Multi-Task LLMs
von: Wu, Runzhe, et al.
Veröffentlicht: (2025)
von: Wu, Runzhe, et al.
Veröffentlicht: (2025)
Post-Training with Policy Gradients: Optimality and the Base Model Barrier
von: Mousavi-Hosseini, Alireza, et al.
Veröffentlicht: (2026)
von: Mousavi-Hosseini, Alireza, et al.
Veröffentlicht: (2026)
Training Data Selection with Gradient Orthogonality for Efficient Domain Adaptation
von: Zhang, Xiyang, et al.
Veröffentlicht: (2026)
von: Zhang, Xiyang, et al.
Veröffentlicht: (2026)
AnyBipe: An End-to-End Framework for Training and Deploying Bipedal Robots Guided by Large Language Models
von: Yao, Yifei, et al.
Veröffentlicht: (2024)
von: Yao, Yifei, et al.
Veröffentlicht: (2024)
Accelerating RLHF Training with Reward Variance Increase
von: Yang, Zonglin, et al.
Veröffentlicht: (2025)
von: Yang, Zonglin, et al.
Veröffentlicht: (2025)
When and Why Adversarial Training Improves PINNs: A Neural Tangent Kernel Perspective
von: Cao, Yuan-dong, et al.
Veröffentlicht: (2026)
von: Cao, Yuan-dong, et al.
Veröffentlicht: (2026)
FullCert: Deterministic End-to-End Certification for Training and Inference of Neural Networks
von: Lorenz, Tobias, et al.
Veröffentlicht: (2024)
von: Lorenz, Tobias, et al.
Veröffentlicht: (2024)
TSPulse: Tiny Pre-Trained Models with Disentangled Representations for Rapid Time-Series Analysis
von: Ekambaram, Vijay, et al.
Veröffentlicht: (2025)
von: Ekambaram, Vijay, et al.
Veröffentlicht: (2025)
Training Large Language Models to Reason via EM Policy Gradient
von: Xu, Tianbing
Veröffentlicht: (2025)
von: Xu, Tianbing
Veröffentlicht: (2025)
Understanding Gradient Boosting Classifier: Training, Prediction, and the Role of $γ_j$
von: Chen, Hung-Hsuan
Veröffentlicht: (2024)
von: Chen, Hung-Hsuan
Veröffentlicht: (2024)
Gradient Weight-normalized Low-rank Projection for Efficient LLM Training
von: Huang, Jia-Hong, et al.
Veröffentlicht: (2024)
von: Huang, Jia-Hong, et al.
Veröffentlicht: (2024)
Dynamic Tokenization via Reinforcement Patching: End-to-end Training and Zero-shot Transfer
von: Wu, Yulun, et al.
Veröffentlicht: (2026)
von: Wu, Yulun, et al.
Veröffentlicht: (2026)
Training and Simulation of Quadrupedal Robot in Adaptive Stair Climbing for Indoor Firefighting: An End-to-End Reinforcement Learning Approach
von: Huang, Baixiao, et al.
Veröffentlicht: (2026)
von: Huang, Baixiao, et al.
Veröffentlicht: (2026)
RapidGNN: Energy and Communication-Efficient Distributed Training on Large-Scale Graph Neural Networks
von: Niam, Arefin, et al.
Veröffentlicht: (2025)
von: Niam, Arefin, et al.
Veröffentlicht: (2025)
Gradient-Informed Temporal Sampling Improves Rollout Accuracy in PDE Surrogate Training
von: Wang, Wenshuo, et al.
Veröffentlicht: (2026)
von: Wang, Wenshuo, et al.
Veröffentlicht: (2026)
GAC: Stabilizing Asynchronous RL Training for LLMs via Gradient Alignment Control
von: Xu, Haofeng, et al.
Veröffentlicht: (2026)
von: Xu, Haofeng, et al.
Veröffentlicht: (2026)
Beyond the Mean: Fisher-Orthogonal Projection for Natural Gradient Descent in Large Batch Training
von: Lu, Yishun, et al.
Veröffentlicht: (2025)
von: Lu, Yishun, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
ScheduleFree+: Scaling Learning-Rate-Free & Schedule-Free Learning to Large Language Models
von: Defazio, Aaron
Veröffentlicht: (2026) -
Prodigy: An Expeditiously Adaptive Parameter-Free Learner
von: Mishchenko, Konstantin, et al.
Veröffentlicht: (2023) -
Smoothing DiLoCo with Primal Averaging for Faster Training of LLMs
von: Defazio, Aaron, et al.
Veröffentlicht: (2025) -
Optimal Linear Decay Learning Rate Schedules and Further Refinements
von: Defazio, Aaron, et al.
Veröffentlicht: (2023) -
Purifying Shampoo: Investigating Shampoo's Heuristics by Decomposing its Preconditioner
von: Eschenhagen, Runa, et al.
Veröffentlicht: (2025)