Optimal Linear Decay Learning Rate Schedules and Further Refinements
Fuente:
arXiv
Saved in:
| Main Authors: | Defazio, Aaron, Cutkosky, Ashok, Mehta, Harsh, Mishchenko, Konstantin |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Road Less Scheduled
by: Defazio, Aaron, et al.
Published: (2024)
by: Defazio, Aaron, et al.
Published: (2024)
Prodigy: An Expeditiously Adaptive Parameter-Free Learner
by: Mishchenko, Konstantin, et al.
Published: (2023)
by: Mishchenko, Konstantin, et al.
Published: (2023)
ScheduleFree+: Scaling Learning-Rate-Free & Schedule-Free Learning to Large Language Models
by: Defazio, Aaron
Published: (2026)
by: Defazio, Aaron
Published: (2026)
Smoothing DiLoCo with Primal Averaging for Faster Training of LLMs
by: Defazio, Aaron, et al.
Published: (2025)
by: Defazio, Aaron, et al.
Published: (2025)
Why Gradients Rapidly Increase Near the End of Training
by: Defazio, Aaron
Published: (2025)
by: Defazio, Aaron
Published: (2025)
Optimal Stochastic Non-smooth Non-convex Optimization through Online-to-Non-convex Conversion
by: Cutkosky, Ashok, et al.
Published: (2023)
by: Cutkosky, Ashok, et al.
Published: (2023)
Training Dynamics of the Cooldown Stage in Warmup-Stable-Decay Learning Rate Scheduler
by: Dremov, Aleksandr, et al.
Published: (2025)
by: Dremov, Aleksandr, et al.
Published: (2025)
Optimal Decay Spectra for Linear Recurrences
by: Cao, Yang
Published: (2026)
by: Cao, Yang
Published: (2026)
Cyclical Log Annealing as a Learning Rate Scheduler
by: Naveen, Philip
Published: (2024)
by: Naveen, Philip
Published: (2024)
Online Linear Regression in Dynamic Environments via Discounting
by: Jacobsen, Andrew, et al.
Published: (2024)
by: Jacobsen, Andrew, et al.
Published: (2024)
Theoretical Analysis of Sparse Optimization with Reparameterization, Weight Decay, and Adaptive Learning Rate
by: Xu, Huangyu, et al.
Published: (2026)
by: Xu, Huangyu, et al.
Published: (2026)
Purifying Shampoo: Investigating Shampoo's Heuristics by Decomposing its Preconditioner
by: Eschenhagen, Runa, et al.
Published: (2025)
by: Eschenhagen, Runa, et al.
Published: (2025)
Adaptive Memory Decay for Log-Linear Attention
by: Amin, Yaxita, et al.
Published: (2026)
by: Amin, Yaxita, et al.
Published: (2026)
Heterogeneous Learning Rate Scheduling for Neural Architecture Search on Long-Tailed Datasets
by: Tang, Chenxia
Published: (2024)
by: Tang, Chenxia
Published: (2024)
Straight to Zero: Why Linearly Decaying the Learning Rate to Zero Works Best for LLMs
by: Bergsma, Shane, et al.
Published: (2025)
by: Bergsma, Shane, et al.
Published: (2025)
Localized Observation Abstraction Using Piecewise Linear Spatial Decay for Reinforcement Learning in Combat Simulations
by: Black, Scotty, et al.
Published: (2024)
by: Black, Scotty, et al.
Published: (2024)
Power Scheduler: A Batch Size and Token Number Agnostic Learning Rate Scheduler
by: Shen, Yikang, et al.
Published: (2024)
by: Shen, Yikang, et al.
Published: (2024)
Federated Learning with Neural Graphical Models
by: Chajewska, Urszula, et al.
Published: (2023)
by: Chajewska, Urszula, et al.
Published: (2023)
Improving Adaptive Online Learning Using Refined Discretization
by: Zhang, Zhiyu, et al.
Published: (2023)
by: Zhang, Zhiyu, et al.
Published: (2023)
Seesaw: Accelerating Training by Balancing Learning Rate and Batch Size Scheduling
by: Meterez, Alexandru, et al.
Published: (2025)
by: Meterez, Alexandru, et al.
Published: (2025)
Anytime Pretraining: Horizon-Free Learning-Rate Schedules with Weight Averaging
by: Meterez, Alexandru, et al.
Published: (2026)
by: Meterez, Alexandru, et al.
Published: (2026)
Principled Curriculum Learning using Parameter Continuation Methods
by: Pathak, Harsh Nilesh, et al.
Published: (2025)
by: Pathak, Harsh Nilesh, et al.
Published: (2025)
Learning Graph Node Embeddings by Smooth Pair Sampling
by: Kutzkov, Konstantin
Published: (2025)
by: Kutzkov, Konstantin
Published: (2025)
When Are Learning Biases Equivalent? A Unifying Framework for Fairness, Robustness, and Distribution Shift
by: Mehta, Sushant
Published: (2025)
by: Mehta, Sushant
Published: (2025)
Generative Kaleidoscopic Networks
by: Shrivastava, Harsh
Published: (2024)
by: Shrivastava, Harsh
Published: (2024)
Time Transfer: On Optimal Learning Rate and Batch Size In The Infinite Data Limit
by: Filatov, Oleg, et al.
Published: (2024)
by: Filatov, Oleg, et al.
Published: (2024)
Scalable Production Scheduling: Linear Complexity via Unified Homogeneous Graphs
by: Hoss, Jonathan, et al.
Published: (2026)
by: Hoss, Jonathan, et al.
Published: (2026)
SOAR: Self-Correction for Optimal Alignment and Refinement in Diffusion Models
by: Qin, You, et al.
Published: (2026)
by: Qin, You, et al.
Published: (2026)
Feature Selection via GANs (GANFS): Enhancing Machine Learning Models for DDoS Mitigation
by: Patel, Harsh
Published: (2025)
by: Patel, Harsh
Published: (2025)
How Learning Rate Decay Wastes Your Best Data in Curriculum-Based LLM Pretraining
by: Luo, Kairong, et al.
Published: (2025)
by: Luo, Kairong, et al.
Published: (2025)
General framework for online-to-nonconvex conversion: Schedule-free SGD is also effective for nonconvex optimization
by: Ahn, Kwangjun, et al.
Published: (2024)
by: Ahn, Kwangjun, et al.
Published: (2024)
Globally Optimal Hierarchical Reinforcement Learning for Linearly-Solvable Markov Decision Processes
by: Infante, Guillermo, et al.
Published: (2021)
by: Infante, Guillermo, et al.
Published: (2021)
Fully Unconstrained Online Learning
by: Cutkosky, Ashok, et al.
Published: (2024)
by: Cutkosky, Ashok, et al.
Published: (2024)
Linear $Q$-Learning Does Not Diverge in $L^2$: Convergence Rates to a Bounded Set
by: Liu, Xinyu, et al.
Published: (2025)
by: Liu, Xinyu, et al.
Published: (2025)
ExpTest: Automating Learning Rate Searching and Tuning with Insights from Linearized Neural Networks
by: Chaudhry, Zan, et al.
Published: (2024)
by: Chaudhry, Zan, et al.
Published: (2024)
A Multi-Power Law for Loss Curve Prediction Across Learning Rate Schedules
by: Luo, Kairong, et al.
Published: (2025)
by: Luo, Kairong, et al.
Published: (2025)
Scaling Laws and In-Context Learning: A Unified Theoretical Framework
by: Mehta, Sushant, et al.
Published: (2025)
by: Mehta, Sushant, et al.
Published: (2025)
Local Linear Attention: An Optimal Interpolation of Linear and Softmax Attention For Test-Time Regression
by: Zuo, Yifei, et al.
Published: (2025)
by: Zuo, Yifei, et al.
Published: (2025)
Exploring the Performance of Perforated Backpropagation through Further Experiments
by: Brenner, Rorry, et al.
Published: (2025)
by: Brenner, Rorry, et al.
Published: (2025)
Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size
by: Hayou, Soufiane, et al.
Published: (2025)
by: Hayou, Soufiane, et al.
Published: (2025)
Similar Items
-
The Road Less Scheduled
by: Defazio, Aaron, et al.
Published: (2024) -
Prodigy: An Expeditiously Adaptive Parameter-Free Learner
by: Mishchenko, Konstantin, et al.
Published: (2023) -
ScheduleFree+: Scaling Learning-Rate-Free & Schedule-Free Learning to Large Language Models
by: Defazio, Aaron
Published: (2026) -
Smoothing DiLoCo with Primal Averaging for Faster Training of LLMs
by: Defazio, Aaron, et al.
Published: (2025) -
Why Gradients Rapidly Increase Near the End of Training
by: Defazio, Aaron
Published: (2025)