Adan: Adaptive Nesterov Momentum Algorithm for Faster Optimizing Deep Models
Fuente:
arXiv
Saved in:
| Main Authors: | Xie, Xingyu, Zhou, Pan, Li, Huan, Lin, Zhouchen, Yan, Shuicheng |
|---|---|
| Format: | Preprint |
| Published: |
2022
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Accelerated Gradient Tracking over Time-varying Graphs for Decentralized Optimization
by: Li, Huan, et al.
Published: (2021)
by: Li, Huan, et al.
Published: (2021)
EMA-Nesterov: Stabilizing Nesterov's Lookahead for Accelerated Deep Learning Optimization
by: Yau, Chung-Yiu, et al.
Published: (2026)
by: Yau, Chung-Yiu, et al.
Published: (2026)
Optimization Hyper-parameter Laws for Large Language Models
by: Xie, Xingyu, et al.
Published: (2024)
by: Xie, Xingyu, et al.
Published: (2024)
Provable Accelerated Convergence of Nesterov's Momentum for Deep ReLU Neural Networks
by: Liao, Fangshuo, et al.
Published: (2023)
by: Liao, Fangshuo, et al.
Published: (2023)
On the $O(\frac{\sqrt{d}}{K^{1/4}})$ Convergence Rate of AdamW Measured by $\ell_1$ Norm
by: Li, Huan, et al.
Published: (2025)
by: Li, Huan, et al.
Published: (2025)
Convergence Rate Analysis of LION
by: Dong, Yiming, et al.
Published: (2024)
by: Dong, Yiming, et al.
Published: (2024)
Convergence Rate Analysis of the AdamW-Style Shampoo: Unifying One-Sided and Two-Sided Preconditioning
by: Li, Huan, et al.
Published: (2026)
by: Li, Huan, et al.
Published: (2026)
Stability and Generalization for Stochastic Recursive Momentum-based Algorithms for (Strongly-)Convex One to $K$-Level Stochastic Optimizations
by: Pan, Xiaokang, et al.
Published: (2024)
by: Pan, Xiaokang, et al.
Published: (2024)
Nesterov Finds GRAAL: Optimal and Adaptive Gradient Method for Convex Optimization
by: Borodich, Ekaterina, et al.
Published: (2025)
by: Borodich, Ekaterina, et al.
Published: (2025)
Muon with Nesterov Momentum: Heavy-Tailed Noise and (Randomized) Inexact Polar Decomposition
by: Choudhury, Sayantan, et al.
Published: (2026)
by: Choudhury, Sayantan, et al.
Published: (2026)
Faster Adaptive Decentralized Learning Algorithms
by: Huang, Feihu, et al.
Published: (2024)
by: Huang, Feihu, et al.
Published: (2024)
Muon is Provably Faster with Momentum Variance Reduction
by: Qian, Xun, et al.
Published: (2025)
by: Qian, Xun, et al.
Published: (2025)
LoCo: Low-Bit Communication Adaptor for Large-scale Model Training
by: Xie, Xingyu, et al.
Published: (2024)
by: Xie, Xingyu, et al.
Published: (2024)
Shuffling Momentum Gradient Algorithm for Convex Optimization
by: Tran, Trang H., et al.
Published: (2024)
by: Tran, Trang H., et al.
Published: (2024)
Provably Faster Algorithms for Bilevel Optimization via Without-Replacement Sampling
by: Li, Junyi, et al.
Published: (2024)
by: Li, Junyi, et al.
Published: (2024)
Beyond likelihood ratio bias: Nested multi-time-scale stochastic approximation for likelihood-free parameter estimation
by: Li, Zehao, et al.
Published: (2024)
by: Li, Zehao, et al.
Published: (2024)
Adaptive Optimization via Momentum on Variance-Normalized Gradients
by: Patitucci, Francisco, et al.
Published: (2026)
by: Patitucci, Francisco, et al.
Published: (2026)
Faster Gradient-Free Algorithms for Nonsmooth Nonconvex Stochastic Optimization
by: Chen, Lesi, et al.
Published: (2023)
by: Chen, Lesi, et al.
Published: (2023)
Adaptive Moment Estimation Optimization Algorithm Using Projection Gradient for Deep Learning
by: Li, Yongqi, et al.
Published: (2025)
by: Li, Yongqi, et al.
Published: (2025)
Generalized Continuous-Time Models for Nesterov's Accelerated Gradient Methods
by: Park, Chanwoong, et al.
Published: (2024)
by: Park, Chanwoong, et al.
Published: (2024)
Faster Stochastic Algorithms for Minimax Optimization under Polyak--Łojasiewicz Conditions
by: Chen, Lesi, et al.
Published: (2023)
by: Chen, Lesi, et al.
Published: (2023)
Nesterov acceleration in benignly non-convex landscapes
by: Gupta, Kanan, et al.
Published: (2024)
by: Gupta, Kanan, et al.
Published: (2024)
A Concise Lyapunov Analysis of Nesterov's Accelerated Gradient Method
by: Liu, Jun
Published: (2025)
by: Liu, Jun
Published: (2025)
Nesterov Acceleration for Ensemble Kalman Inversion and Variants
by: Vernon, Sydney, et al.
Published: (2025)
by: Vernon, Sydney, et al.
Published: (2025)
On the $O(\frac{\sqrt{d}}{T^{1/4}})$ Convergence Rate of RMSProp and Its Momentum Extension Measured by $\ell_1$ Norm
by: Li, Huan, et al.
Published: (2024)
by: Li, Huan, et al.
Published: (2024)
Stochastic Difference-of-Convex Optimization with Momentum
by: Chayti, El Mahdi, et al.
Published: (2025)
by: Chayti, El Mahdi, et al.
Published: (2025)
Faster Algorithms for User-Level Private Stochastic Convex Optimization
by: Lowy, Andrew, et al.
Published: (2024)
by: Lowy, Andrew, et al.
Published: (2024)
Faster Gradient Methods for Highly-Smooth Stochastic Bilevel Optimization
by: Chen, Lesi, et al.
Published: (2025)
by: Chen, Lesi, et al.
Published: (2025)
Inference of Online Newton Methods with Nesterov's Accelerated Sketching
by: Wang, Haoxuan, et al.
Published: (2026)
by: Wang, Haoxuan, et al.
Published: (2026)
SGD with Adaptive Preconditioning: Unified Analysis and Momentum Acceleration
by: Kovalev, Dmitry
Published: (2025)
by: Kovalev, Dmitry
Published: (2025)
Adaptive Momentum and Nonlinear Damping for Neural Network Training
by: Karoni, Aikaterini, et al.
Published: (2026)
by: Karoni, Aikaterini, et al.
Published: (2026)
Faster Algorithms for Structured Linear and Kernel Support Vector Machines
by: Gu, Yuzhou, et al.
Published: (2023)
by: Gu, Yuzhou, et al.
Published: (2023)
Convergence Rate Analysis of SOAP with Arbitrary Orthogonal Projection Matrices
by: Li, Huan, et al.
Published: (2026)
by: Li, Huan, et al.
Published: (2026)
Provable Acceleration of Nesterov's Accelerated Gradient for Rectangular Matrix Factorization and Linear Neural Networks
by: Xu, Zhenghao, et al.
Published: (2024)
by: Xu, Zhenghao, et al.
Published: (2024)
TiAda: A Time-scale Adaptive Algorithm for Nonconvex Minimax Optimization
by: Li, Xiang, et al.
Published: (2022)
by: Li, Xiang, et al.
Published: (2022)
Nesterov acceleration despite very noisy gradients
by: Gupta, Kanan, et al.
Published: (2023)
by: Gupta, Kanan, et al.
Published: (2023)
Non-convex Stochastic Composite Optimization with Polyak Momentum
by: Gao, Yuan, et al.
Published: (2024)
by: Gao, Yuan, et al.
Published: (2024)
On the Provable Suboptimality of Momentum SGD in Nonstationary Stochastic Optimization
by: Sahu, Sharan, et al.
Published: (2026)
by: Sahu, Sharan, et al.
Published: (2026)
Hindsight-Guided Momentum (HGM) Optimizer: An Approach to Adaptive Learning Rate
by: Sarkar, Krisanu
Published: (2025)
by: Sarkar, Krisanu
Published: (2025)
Adaptive Algorithms with Sharp Convergence Rates for Stochastic Hierarchical Optimization
by: Gong, Xiaochuan, et al.
Published: (2025)
by: Gong, Xiaochuan, et al.
Published: (2025)
Similar Items
-
Accelerated Gradient Tracking over Time-varying Graphs for Decentralized Optimization
by: Li, Huan, et al.
Published: (2021) -
EMA-Nesterov: Stabilizing Nesterov's Lookahead for Accelerated Deep Learning Optimization
by: Yau, Chung-Yiu, et al.
Published: (2026) -
Optimization Hyper-parameter Laws for Large Language Models
by: Xie, Xingyu, et al.
Published: (2024) -
Provable Accelerated Convergence of Nesterov's Momentum for Deep ReLU Neural Networks
by: Liao, Fangshuo, et al.
Published: (2023) -
On the $O(\frac{\sqrt{d}}{K^{1/4}})$ Convergence Rate of AdamW Measured by $\ell_1$ Norm
by: Li, Huan, et al.
Published: (2025)