On the $O(\frac{\sqrt{d}}{K^{1/4}})$ Convergence Rate of AdamW Measured by $\ell_1$ Norm
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Huan, Dong, Yiming, Lin, Zhouchen |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Convergence Rate Analysis of the AdamW-Style Shampoo: Unifying One-Sided and Two-Sided Preconditioning
by: Li, Huan, et al.
Published: (2026)
by: Li, Huan, et al.
Published: (2026)
On the $O(\frac{\sqrt{d}}{T^{1/4}})$ Convergence Rate of RMSProp and Its Momentum Extension Measured by $\ell_1$ Norm
by: Li, Huan, et al.
Published: (2024)
by: Li, Huan, et al.
Published: (2024)
Implicit Bias of AdamW: $\ell_\infty$ Norm Constrained Optimization
by: Xie, Shuo, et al.
Published: (2024)
by: Xie, Shuo, et al.
Published: (2024)
Convergence Rate Analysis of LION
by: Dong, Yiming, et al.
Published: (2024)
by: Dong, Yiming, et al.
Published: (2024)
HomeAdam: Adam and AdamW Algorithms Sometimes Go Home to Obtain Better Provable Generalization
by: Huang, Feihu, et al.
Published: (2026)
by: Huang, Feihu, et al.
Published: (2026)
Optimizer-Induced Mode Connectivity: From AdamW to Muon
by: Zhang, Fangzhao, et al.
Published: (2026)
by: Zhang, Fangzhao, et al.
Published: (2026)
Accelerated Gradient Tracking over Time-varying Graphs for Decentralized Optimization
by: Li, Huan, et al.
Published: (2021)
by: Li, Huan, et al.
Published: (2021)
Convergence Rate Analysis of SOAP with Arbitrary Orthogonal Projection Matrices
by: Li, Huan, et al.
Published: (2026)
by: Li, Huan, et al.
Published: (2026)
Adam-HNAG: A Convergent Reformulation of Adam with Accelerated Rate
by: Yu, Yaxin, et al.
Published: (2026)
by: Yu, Yaxin, et al.
Published: (2026)
On Convergence of Adam for Stochastic Optimization under Relaxed Assumptions
by: Hong, Yusu, et al.
Published: (2024)
by: Hong, Yusu, et al.
Published: (2024)
Clipping Improves Adam-Norm and AdaGrad-Norm when the Noise Is Heavy-Tailed
by: Chezhegov, Savelii, et al.
Published: (2024)
by: Chezhegov, Savelii, et al.
Published: (2024)
Adan: Adaptive Nesterov Momentum Algorithm for Faster Optimizing Deep Models
by: Xie, Xingyu, et al.
Published: (2022)
by: Xie, Xingyu, et al.
Published: (2022)
Beyond likelihood ratio bias: Nested multi-time-scale stochastic approximation for likelihood-free parameter estimation
by: Li, Zehao, et al.
Published: (2024)
by: Li, Zehao, et al.
Published: (2024)
Adam Converges Without Any Modification On Update Rules
by: Zhang, Yushun, et al.
Published: (2026)
by: Zhang, Yushun, et al.
Published: (2026)
Convergence rates for the Adam optimizer
by: Dereich, Steffen, et al.
Published: (2024)
by: Dereich, Steffen, et al.
Published: (2024)
Adam-SHANG: A Convergent Adam-Type Method for Stochastic Smooth Convex Optimization
by: Yu, Yaxin, et al.
Published: (2026)
by: Yu, Yaxin, et al.
Published: (2026)
A Comprehensive Framework for Analyzing the Convergence of Adam: Bridging the Gap with SGD
by: Jin, Ruinan, et al.
Published: (2024)
by: Jin, Ruinan, et al.
Published: (2024)
Adam-family Methods for Nonsmooth Optimization with Convergence Guarantees
by: Xiao, Nachuan, et al.
Published: (2023)
by: Xiao, Nachuan, et al.
Published: (2023)
De-singularity Subgradient for the $q$-th-Powered $\ell_p$-Norm Weber Location Problem
by: Lai, Zhao-Rong, et al.
Published: (2024)
by: Lai, Zhao-Rong, et al.
Published: (2024)
Convergence of Steepest Descent and Adam under Non-Uniform Smoothness
by: Vaswani, Sharan, et al.
Published: (2026)
by: Vaswani, Sharan, et al.
Published: (2026)
On the Convergence of Adam-Type Algorithm for Bilevel Optimization under Unbounded Smoothness
by: Gong, Xiaochuan, et al.
Published: (2025)
by: Gong, Xiaochuan, et al.
Published: (2025)
Reusing Historical Trajectories in Natural Policy Gradient via Importance Sampling: Convergence and Convergence Rate
by: Lin, Yifan, et al.
Published: (2024)
by: Lin, Yifan, et al.
Published: (2024)
On the Convergence of Adam under Non-uniform Smoothness: Separability from SGDM and Beyond
by: Wang, Bohan, et al.
Published: (2024)
by: Wang, Bohan, et al.
Published: (2024)
A Determinantal Approach to a Sharp $\ell^1-\ell^\infty-\ell^2$ Norm Inequality
by: Benitez, Jose Antonio Lara
Published: (2026)
by: Benitez, Jose Antonio Lara
Published: (2026)
Convergence Guarantees for RMSProp and Adam in Generalized-smooth Non-convex Optimization with Affine Noise Variance
by: Zhang, Qi, et al.
Published: (2024)
by: Zhang, Qi, et al.
Published: (2024)
Learning Spatially Adaptive $\ell_1$-Norms Weights for Convolutional Synthesis Regularization
by: Kofler, Andreas, et al.
Published: (2025)
by: Kofler, Andreas, et al.
Published: (2025)
A Theoretical and Empirical Study on the Convergence of Adam with an "Exact" Constant Step Size in Non-Convex Settings
by: Mazumder, Alokendu, et al.
Published: (2023)
by: Mazumder, Alokendu, et al.
Published: (2023)
Sparse Deep Learning Models with the $\ell_1$ Regularization
by: Shen, Lixin, et al.
Published: (2024)
by: Shen, Lixin, et al.
Published: (2024)
MAP Estimation with Denoisers: Convergence Rates and Guarantees
by: Pesme, Scott, et al.
Published: (2025)
by: Pesme, Scott, et al.
Published: (2025)
Robust Sublinear Convergence Rates for Iterative Bregman Projections
by: Peyré, Gabriel
Published: (2026)
by: Peyré, Gabriel
Published: (2026)
Limits of Convergence-Rate Control for Open-Weight Safety
by: Rosati, Domenic, et al.
Published: (2026)
by: Rosati, Domenic, et al.
Published: (2026)
Improved Convergence Rates of Muon Optimizer for Nonconvex Optimization
by: Nagashima, Shuntaro, et al.
Published: (2026)
by: Nagashima, Shuntaro, et al.
Published: (2026)
Implicit Bias and Fast Convergence Rates for Self-attention
by: Vasudeva, Bhavya, et al.
Published: (2024)
by: Vasudeva, Bhavya, et al.
Published: (2024)
Convergence Rate of the Last Iterate of Stochastic Proximal Algorithms
by: Vaidyan, Kevin Kurian Thomas, et al.
Published: (2026)
by: Vaidyan, Kevin Kurian Thomas, et al.
Published: (2026)
Open Problem: Anytime Convergence Rate of Gradient Descent
by: Kornowski, Guy, et al.
Published: (2024)
by: Kornowski, Guy, et al.
Published: (2024)
Incremental Gauss--Newton Methods with Superlinear Convergence Rates
by: Zhou, Zhiling, et al.
Published: (2024)
by: Zhou, Zhiling, et al.
Published: (2024)
Nonsmooth Implicit Differentiation: Deterministic and Stochastic Convergence Rates
by: Grazzi, Riccardo, et al.
Published: (2024)
by: Grazzi, Riccardo, et al.
Published: (2024)
Adam on Local Time: Addressing Nonstationarity in RL with Relative Adam Timesteps
by: Ellis, Benjamin, et al.
Published: (2024)
by: Ellis, Benjamin, et al.
Published: (2024)
How to Set $β_1, β_2$ in Adam: An Online Learning Perspective
by: Nguyen, Quan
Published: (2025)
by: Nguyen, Quan
Published: (2025)
Understanding Adam Optimizer via Online Learning of Updates: Adam is FTRL in Disguise
by: Ahn, Kwangjun, et al.
Published: (2024)
by: Ahn, Kwangjun, et al.
Published: (2024)
Similar Items
-
Convergence Rate Analysis of the AdamW-Style Shampoo: Unifying One-Sided and Two-Sided Preconditioning
by: Li, Huan, et al.
Published: (2026) -
On the $O(\frac{\sqrt{d}}{T^{1/4}})$ Convergence Rate of RMSProp and Its Momentum Extension Measured by $\ell_1$ Norm
by: Li, Huan, et al.
Published: (2024) -
Implicit Bias of AdamW: $\ell_\infty$ Norm Constrained Optimization
by: Xie, Shuo, et al.
Published: (2024) -
Convergence Rate Analysis of LION
by: Dong, Yiming, et al.
Published: (2024) -
HomeAdam: Adam and AdamW Algorithms Sometimes Go Home to Obtain Better Provable Generalization
by: Huang, Feihu, et al.
Published: (2026)