Straight to Zero: Why Linearly Decaying the Learning Rate to Zero Works Best for LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Bergsma, Shane, Dey, Nolan, Gosal, Gurpreet, Gray, Gavia, Soboleva, Daria, Hestness, Joel |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Power Lines: Scaling Laws for Weight Decay and Batch Size in LLM Pre-training
by: Bergsma, Shane, et al.
Published: (2025)
by: Bergsma, Shane, et al.
Published: (2025)
Predicting Training Re-evaluation Curves Enables Effective Data Curriculums for LLMs
by: Bergsma, Shane, et al.
Published: (2025)
by: Bergsma, Shane, et al.
Published: (2025)
A Review on Zeroing Neural Networks
by: Jiang, Chengze, et al.
Published: (2025)
by: Jiang, Chengze, et al.
Published: (2025)
Scaling with Collapse: Efficient and Predictable Training of LLM Families
by: Bergsma, Shane, et al.
Published: (2025)
by: Bergsma, Shane, et al.
Published: (2025)
Neural Optimizer Equation, Decay Function, and Learning Rate Schedule Joint Evolution
by: Morgan, Brandon, et al.
Published: (2024)
by: Morgan, Brandon, et al.
Published: (2024)
Matrix Domination: Convergence of a Genetic Algorithm Metaheuristic with the Wisdom of Crowds to Solve the NP-Complete Problem
by: Strachan, Shane Storm
Published: (2023)
by: Strachan, Shane Storm
Published: (2023)
Plug-and-Play Homeostatic Spark: Zero-Cost Acceleration for SNN Training Across Paradigms
by: Chen, Rui, et al.
Published: (2025)
by: Chen, Rui, et al.
Published: (2025)
Zero-Shot Document-Level Biomedical Relation Extraction via Scenario-based Prompt Design in Two-Stage with LLM
by: Zhao, Lei, et al.
Published: (2025)
by: Zhao, Lei, et al.
Published: (2025)
Connectome-Guided Automatic Learning Rates for Deep Networks
by: He, Peilin, et al.
Published: (2025)
by: He, Peilin, et al.
Published: (2025)
Learning to Forget: Continual Learning with Adaptive Weight Decay
by: Ramesh, Aditya A., et al.
Published: (2026)
by: Ramesh, Aditya A., et al.
Published: (2026)
EvoGM: Learning to Merge LLMs via Evolutionary Generative Optimization
by: Jiang, Tao, et al.
Published: (2026)
by: Jiang, Tao, et al.
Published: (2026)
Pretrained Optimization Model for Zero-Shot Black Box Optimization
by: Li, Xiaobin, et al.
Published: (2024)
by: Li, Xiaobin, et al.
Published: (2024)
Why Prompt Optimization Works, and Why It Sometimes Doesn't: A Causal-Inspired Edit-Level Analysis
by: Gong, Shuzhi, et al.
Published: (2026)
by: Gong, Shuzhi, et al.
Published: (2026)
UniHPF : Universal Healthcare Predictive Framework with Zero Domain Knowledge
by: Hur, Kyunghoon, et al.
Published: (2022)
by: Hur, Kyunghoon, et al.
Published: (2022)
Evolutionary Generalized Zero-Shot Learning
by: Chen, Dubing, et al.
Published: (2022)
by: Chen, Dubing, et al.
Published: (2022)
Effective Adaptive Mutation Rates for Program Synthesis
by: Ni, Andrew, et al.
Published: (2024)
by: Ni, Andrew, et al.
Published: (2024)
GreenMachine: Automatic Design of Zero-Cost Proxies for Energy-Efficient NAS
by: Cortês, Gabriel, et al.
Published: (2024)
by: Cortês, Gabriel, et al.
Published: (2024)
A Predefined-Time Convergent and Noise-Tolerant Zeroing Neural Network Model for Time Variant Quadratic Programming With Application to Robot Motion Planning
by: Yang, Yi, et al.
Published: (2024)
by: Yang, Yi, et al.
Published: (2024)
A Theoretical Perspective on Why Stochastic Population Update Needs an Archive in Evolutionary Multi-objective Optimization
by: Ren, Shengjie, et al.
Published: (2025)
by: Ren, Shengjie, et al.
Published: (2025)
A Flexible Evolutionary Algorithm With Dynamic Mutation Rate Archive
by: Krejca, Martin S., et al.
Published: (2024)
by: Krejca, Martin S., et al.
Published: (2024)
All Constant Mutation Rates for the $(1+1)$ Evolutionary Algorithm
by: Kelley, Andrew James
Published: (2026)
by: Kelley, Andrew James
Published: (2026)
PRIMETIME : Limits of LLMs in Temporal Primitives
by: Gaere, Edward, et al.
Published: (2025)
by: Gaere, Edward, et al.
Published: (2025)
Relation Reasoning with LLMs in Expensive Optimization
by: Lu, Ye, et al.
Published: (2026)
by: Lu, Ye, et al.
Published: (2026)
All Mutation Rates $c/n$ for the $(1+1)$ Evolutionary Algorithm
by: Kelley, Andrew James
Published: (2026)
by: Kelley, Andrew James
Published: (2026)
Hidden Traveling Waves bind Working Memory Variables in Recurrent Neural Networks
by: Karuvally, Arjun, et al.
Published: (2024)
by: Karuvally, Arjun, et al.
Published: (2024)
GP and LLMs for Program Synthesis: No Clear Winners
by: Hernandez, Jose Guadalupe, et al.
Published: (2025)
by: Hernandez, Jose Guadalupe, et al.
Published: (2025)
Analysing Rescaling, Discretization, and Linearization in RNNs for Neural System Modelling
by: Caruso, Mariano, et al.
Published: (2023)
by: Caruso, Mariano, et al.
Published: (2023)
Runtime Analysis of Evolutionary Diversity Optimization on the Multi-objective (LeadingOnes, TrailingZeros) Problem
by: Antipov, Denis, et al.
Published: (2024)
by: Antipov, Denis, et al.
Published: (2024)
Tight Runtime Bounds for Static Unary Unbiased Evolutionary Algorithms on Linear Functions
by: Doerr, Carola, et al.
Published: (2023)
by: Doerr, Carola, et al.
Published: (2023)
Energy Decay Network (EDeN)
by: Shelley, Jamie Nicholas, et al.
Published: (2021)
by: Shelley, Jamie Nicholas, et al.
Published: (2021)
Approximation of a Pareto Set Segment Using a Linear Model with Sharing Variables
by: Guo, Ping, et al.
Published: (2024)
by: Guo, Ping, et al.
Published: (2024)
More than MACs: Exploring the Role of Neuromorphic Engineering in the Age of LLMs
by: Olin-Ammentorp, Wilkie
Published: (2025)
by: Olin-Ammentorp, Wilkie
Published: (2025)
Improving a Parallel C++ Intel AVX-512 SIMD Linear Genetic Programming Interpreter
by: Langdon, William B.
Published: (2025)
by: Langdon, William B.
Published: (2025)
Optimal Distribution of Solutions for Crowding Distance on Linear Pareto Fronts of Two-Objective Optimization Problems
by: Ishibuchi, Hisao, et al.
Published: (2025)
by: Ishibuchi, Hisao, et al.
Published: (2025)
Random-Key Optimizer and Linearization for the Quadratic Multiple Constraints Variable-Sized Bin Packing Problem
by: Santos, Natalia A., et al.
Published: (2025)
by: Santos, Natalia A., et al.
Published: (2025)
CMA-ES with Learning Rate Adaptation
by: Nomura, Masahiro, et al.
Published: (2024)
by: Nomura, Masahiro, et al.
Published: (2024)
Symbolically Regressing Fish Biomass Spectral Data: A Linear Genetic Programming Method with Tunable Primitives
by: Huang, Zhixing, et al.
Published: (2025)
by: Huang, Zhixing, et al.
Published: (2025)
Pruner-Zero: Evolving Symbolic Pruning Metric from scratch for Large Language Models
by: Dong, Peijie, et al.
Published: (2024)
by: Dong, Peijie, et al.
Published: (2024)
Exploring and Learning Structure: Active Inference Approach in Navigational Agents
by: de Tinguy, Daria, et al.
Published: (2024)
by: de Tinguy, Daria, et al.
Published: (2024)
Zero-shot Quantum Neural Architecture Search
by: Dao, Tung, et al.
Published: (2026)
by: Dao, Tung, et al.
Published: (2026)
Similar Items
-
Power Lines: Scaling Laws for Weight Decay and Batch Size in LLM Pre-training
by: Bergsma, Shane, et al.
Published: (2025) -
Predicting Training Re-evaluation Curves Enables Effective Data Curriculums for LLMs
by: Bergsma, Shane, et al.
Published: (2025) -
A Review on Zeroing Neural Networks
by: Jiang, Chengze, et al.
Published: (2025) -
Scaling with Collapse: Efficient and Predictable Training of LLM Families
by: Bergsma, Shane, et al.
Published: (2025) -
Neural Optimizer Equation, Decay Function, and Learning Rate Schedule Joint Evolution
by: Morgan, Brandon, et al.
Published: (2024)