Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Yuxing, Ge, Yuze, Pan, Rui, Kang, An, Zhang, Tong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Accelerated Convergence of Stochastic Heavy Ball Method under Anisotropic Gradient Noise
by: Pan, Rui, et al.
Published: (2023)
by: Pan, Rui, et al.
Published: (2023)
AdaGrad under Anisotropic Smoothness
by: Liu, Yuxing, et al.
Published: (2024)
by: Liu, Yuxing, et al.
Published: (2024)
ASGO: Adaptive Structured Gradient Optimization
by: An, Kang, et al.
Published: (2025)
by: An, Kang, et al.
Published: (2025)
StoSignSGD: Unbiased Structural Stochasticity Fixes SignSGD for Training Large Language Models
by: Yu, Dingzhi, et al.
Published: (2026)
by: Yu, Dingzhi, et al.
Published: (2026)
Finite-Time Decoupled Convergence in Nonlinear Two-Time-Scale Stochastic Approximation
by: Han, Yuze, et al.
Published: (2024)
by: Han, Yuze, et al.
Published: (2024)
Convergence Rate Analysis of LION
by: Dong, Yiming, et al.
Published: (2024)
by: Dong, Yiming, et al.
Published: (2024)
SOREL: A Stochastic Algorithm for Spectral Risks Minimization
by: Ge, Yuze, et al.
Published: (2024)
by: Ge, Yuze, et al.
Published: (2024)
Unbiased Gradient Low-Rank Projection
by: Pan, Rui, et al.
Published: (2025)
by: Pan, Rui, et al.
Published: (2025)
Adam-HNAG: A Convergent Reformulation of Adam with Accelerated Rate
by: Yu, Yaxin, et al.
Published: (2026)
by: Yu, Yaxin, et al.
Published: (2026)
Adaptive SGD with Line-Search and Polyak Stepsizes: Nonconvex Convergence and Accelerated Rates
by: Wu, Haotian
Published: (2025)
by: Wu, Haotian
Published: (2025)
Almost Sure Convergence Rates and Concentration of Stochastic Approximation and Reinforcement Learning with Markovian Noise
by: Qian, Xiaochi, et al.
Published: (2024)
by: Qian, Xiaochi, et al.
Published: (2024)
Almost Sure Convergence Rates of Stochastic Approximation and Reinforcement Learning via a Poisson-Moreau Drift
by: Liu, Xinyu, et al.
Published: (2026)
by: Liu, Xinyu, et al.
Published: (2026)
Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less
by: Liu, Yuxing, et al.
Published: (2026)
by: Liu, Yuxing, et al.
Published: (2026)
Incremental Gauss--Newton Methods with Superlinear Convergence Rates
by: Zhou, Zhiling, et al.
Published: (2024)
by: Zhou, Zhiling, et al.
Published: (2024)
Adaptive Algorithms with Sharp Convergence Rates for Stochastic Hierarchical Optimization
by: Gong, Xiaochuan, et al.
Published: (2025)
by: Gong, Xiaochuan, et al.
Published: (2025)
Incremental Quasi-Newton Methods with Faster Superlinear Convergence Rates
by: Liu, Zhuanghua, et al.
Published: (2024)
by: Liu, Zhuanghua, et al.
Published: (2024)
Reevaluating Theoretical Analysis Methods for Optimization in Deep Learning
by: Tran, Hoang, et al.
Published: (2024)
by: Tran, Hoang, et al.
Published: (2024)
MAP Estimation with Denoisers: Convergence Rates and Guarantees
by: Pesme, Scott, et al.
Published: (2025)
by: Pesme, Scott, et al.
Published: (2025)
Efficient Sign-Based Optimization: Accelerating Convergence via Variance Reduction
by: Jiang, Wei, et al.
Published: (2024)
by: Jiang, Wei, et al.
Published: (2024)
A Minimax-MDP Framework with Future-imposed Conditions for Learning-augmented Problems
by: Chen, Xin, et al.
Published: (2025)
by: Chen, Xin, et al.
Published: (2025)
Understanding Outer Optimizers in Local SGD: Learning Rates, Momentum, and Acceleration
by: Khaled, Ahmed, et al.
Published: (2025)
by: Khaled, Ahmed, et al.
Published: (2025)
A Systems-Theoretic View on the Convergence of Algorithms under Disturbances
by: Er, Guner Dilsad, et al.
Published: (2025)
by: Er, Guner Dilsad, et al.
Published: (2025)
AdaSwitch: An Adaptive Switching Meta-Algorithm for Learning-Augmented Bounded-Influence Problems
by: Chen, Xi, et al.
Published: (2025)
by: Chen, Xi, et al.
Published: (2025)
Robust Sublinear Convergence Rates for Iterative Bregman Projections
by: Peyré, Gabriel
Published: (2026)
by: Peyré, Gabriel
Published: (2026)
Limits of Convergence-Rate Control for Open-Weight Safety
by: Rosati, Domenic, et al.
Published: (2026)
by: Rosati, Domenic, et al.
Published: (2026)
Improved Convergence Rates of Muon Optimizer for Nonconvex Optimization
by: Nagashima, Shuntaro, et al.
Published: (2026)
by: Nagashima, Shuntaro, et al.
Published: (2026)
Implicit Bias and Fast Convergence Rates for Self-attention
by: Vasudeva, Bhavya, et al.
Published: (2024)
by: Vasudeva, Bhavya, et al.
Published: (2024)
Convergence Rate of the Last Iterate of Stochastic Proximal Algorithms
by: Vaidyan, Kevin Kurian Thomas, et al.
Published: (2026)
by: Vaidyan, Kevin Kurian Thomas, et al.
Published: (2026)
Open Problem: Anytime Convergence Rate of Gradient Descent
by: Kornowski, Guy, et al.
Published: (2024)
by: Kornowski, Guy, et al.
Published: (2024)
Nonsmooth Implicit Differentiation: Deterministic and Stochastic Convergence Rates
by: Grazzi, Riccardo, et al.
Published: (2024)
by: Grazzi, Riccardo, et al.
Published: (2024)
Non-Parametric Learning of Stochastic Differential Equations with Non-asymptotic Fast Rates of Convergence
by: Bonalli, Riccardo, et al.
Published: (2023)
by: Bonalli, Riccardo, et al.
Published: (2023)
Reusing Historical Trajectories in Natural Policy Gradient via Importance Sampling: Convergence and Convergence Rate
by: Lin, Yifan, et al.
Published: (2024)
by: Lin, Yifan, et al.
Published: (2024)
Convergence Rate Analysis of the AdamW-Style Shampoo: Unifying One-Sided and Two-Sided Preconditioning
by: Li, Huan, et al.
Published: (2026)
by: Li, Huan, et al.
Published: (2026)
Optimistic Online-to-Batch Conversions for Accelerated Convergence and Universality
by: Yan, Yu-Hu, et al.
Published: (2025)
by: Yan, Yu-Hu, et al.
Published: (2025)
Convergence of Sharpness-Aware Minimization Algorithms using Increasing Batch Size and Decaying Learning Rate
by: Harada, Hinata, et al.
Published: (2024)
by: Harada, Hinata, et al.
Published: (2024)
Increasing Both Batch Size and Learning Rate Accelerates Stochastic Gradient Descent
by: Umeda, Hikaru, et al.
Published: (2024)
by: Umeda, Hikaru, et al.
Published: (2024)
Convergence Rate in Nonlinear Two-Time-Scale Stochastic Approximation with State (Time)-Dependence
by: Chen, Zixi, et al.
Published: (2025)
by: Chen, Zixi, et al.
Published: (2025)
Sharper Convergence Rates for Nonconvex Optimisation via Reduction Mappings
by: Markou, Evan, et al.
Published: (2025)
by: Markou, Evan, et al.
Published: (2025)
On the Complexity of Finite-Sum Smooth Optimization under the Polyak-Łojasiewicz Condition
by: Bai, Yunyan, et al.
Published: (2024)
by: Bai, Yunyan, et al.
Published: (2024)
Faster Convergence of Stochastic Accelerated Gradient Descent under Interpolation
by: Mishkin, Aaron, et al.
Published: (2024)
by: Mishkin, Aaron, et al.
Published: (2024)
Similar Items
-
Accelerated Convergence of Stochastic Heavy Ball Method under Anisotropic Gradient Noise
by: Pan, Rui, et al.
Published: (2023) -
AdaGrad under Anisotropic Smoothness
by: Liu, Yuxing, et al.
Published: (2024) -
ASGO: Adaptive Structured Gradient Optimization
by: An, Kang, et al.
Published: (2025) -
StoSignSGD: Unbiased Structural Stochasticity Fixes SignSGD for Training Large Language Models
by: Yu, Dingzhi, et al.
Published: (2026) -
Finite-Time Decoupled Convergence in Nonlinear Two-Time-Scale Stochastic Approximation
by: Han, Yuze, et al.
Published: (2024)