A Comprehensive Framework for Analyzing the Convergence of Adam: Bridging the Gap with SGD
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Jin, Ruinan, Li, Xiao, Yu, Yaoliang, Wang, Baoxiang |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
The Rich and the Simple: On the Implicit Bias of Adam and SGD
par: Vasudeva, Bhavya, et autres
Publié: (2025)
par: Vasudeva, Bhavya, et autres
Publié: (2025)
Adam-HNAG: A Convergent Reformulation of Adam with Accelerated Rate
par: Yu, Yaxin, et autres
Publié: (2026)
par: Yu, Yaxin, et autres
Publié: (2026)
Adam-SHANG: A Convergent Adam-Type Method for Stochastic Smooth Convex Optimization
par: Yu, Yaxin, et autres
Publié: (2026)
par: Yu, Yaxin, et autres
Publié: (2026)
Adam-family Methods for Nonsmooth Optimization with Convergence Guarantees
par: Xiao, Nachuan, et autres
Publié: (2023)
par: Xiao, Nachuan, et autres
Publié: (2023)
Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling
par: Srećković, Teodora, et autres
Publié: (2025)
par: Srećković, Teodora, et autres
Publié: (2025)
A Unified Framework for Gradient Aggregation in Multi-Objective Optimization
par: Hu, Zeou, et autres
Publié: (2026)
par: Hu, Zeou, et autres
Publié: (2026)
Stability and convergence analysis of AdaGrad for non-convex optimization via novel stopping time-based techniques
par: Jin, Ruinan, et autres
Publié: (2024)
par: Jin, Ruinan, et autres
Publié: (2024)
Faster Convergence of Local SGD for Over-Parameterized Models
par: Qin, Tiancheng, et autres
Publié: (2022)
par: Qin, Tiancheng, et autres
Publié: (2022)
Global Convergence of SGD On Two Layer Neural Nets
par: Gopalani, Pulkit, et autres
Publié: (2022)
par: Gopalani, Pulkit, et autres
Publié: (2022)
From PowerSGD to PowerSGD+: Low-Rank Gradient Compression for Distributed Optimization with Convergence Guarantees
par: Xie, Shengping, et autres
Publié: (2025)
par: Xie, Shengping, et autres
Publié: (2025)
Adam Converges Without Any Modification On Update Rules
par: Zhang, Yushun, et autres
Publié: (2026)
par: Zhang, Yushun, et autres
Publié: (2026)
Fast Last-Iterate Convergence of SGD in the Smooth Interpolation Regime
par: Attia, Amit, et autres
Publié: (2025)
par: Attia, Amit, et autres
Publié: (2025)
Convergence rates for the Adam optimizer
par: Dereich, Steffen, et autres
Publié: (2024)
par: Dereich, Steffen, et autres
Publié: (2024)
Convergence of SGD with momentum in the nonconvex case: A time window-based analysis
par: Qiu, Junwen, et autres
Publié: (2024)
par: Qiu, Junwen, et autres
Publié: (2024)
VAMO: Efficient Zeroth-Order Variance Reduction for SGD with Faster Convergence
par: Chen, Jiahe, et autres
Publié: (2025)
par: Chen, Jiahe, et autres
Publié: (2025)
Global Convergence of SGD For Logistic Loss on Two Layer Neural Nets
par: Gopalani, Pulkit, et autres
Publié: (2023)
par: Gopalani, Pulkit, et autres
Publié: (2023)
On the Convergence of DP-SGD with Adaptive Clipping
par: Shulgin, Egor, et autres
Publié: (2024)
par: Shulgin, Egor, et autres
Publié: (2024)
On Convergence of Adam for Stochastic Optimization under Relaxed Assumptions
par: Hong, Yusu, et autres
Publié: (2024)
par: Hong, Yusu, et autres
Publié: (2024)
Differentially Private Clipped-SGD: High-Probability Convergence with Arbitrary Clipping Level
par: Khah, Saleh Vatan, et autres
Publié: (2025)
par: Khah, Saleh Vatan, et autres
Publié: (2025)
Adaptive SGD with Line-Search and Polyak Stepsizes: Nonconvex Convergence and Accelerated Rates
par: Wu, Haotian
Publié: (2025)
par: Wu, Haotian
Publié: (2025)
Convergence and concentration properties of constant step-size SGD through Markov chains
par: Merad, Ibrahim, et autres
Publié: (2023)
par: Merad, Ibrahim, et autres
Publié: (2023)
Convergence of Steepest Descent and Adam under Non-Uniform Smoothness
par: Vaswani, Sharan, et autres
Publié: (2026)
par: Vaswani, Sharan, et autres
Publié: (2026)
Diagonalisation SGD: Fast & Convergent SGD for Non-Differentiable Models via Reparameterisation and Smoothing
par: Wagner, Dominik, et autres
Publié: (2024)
par: Wagner, Dominik, et autres
Publié: (2024)
Convergence of SGD for Training Neural Networks with Sliced Wasserstein Losses
par: Tanguy, Eloi
Publié: (2023)
par: Tanguy, Eloi
Publié: (2023)
High-Probability Convergence Guarantees of Decentralized SGD
par: Armacki, Aleksandar, et autres
Publié: (2025)
par: Armacki, Aleksandar, et autres
Publié: (2025)
On the Convergence of Adam under Non-uniform Smoothness: Separability from SGDM and Beyond
par: Wang, Bohan, et autres
Publié: (2024)
par: Wang, Bohan, et autres
Publié: (2024)
SGD at the Edge of Stability: The Stochastic Sharpness Gap
par: Liao, Fangshuo, et autres
Publié: (2026)
par: Liao, Fangshuo, et autres
Publié: (2026)
Does Worst-Performing Agent Lead the Pack? Analyzing Agent Dynamics in Unified Distributed SGD
par: Hu, Jie, et autres
Publié: (2024)
par: Hu, Jie, et autres
Publié: (2024)
On the Convergence of Adam-Type Algorithm for Bilevel Optimization under Unbounded Smoothness
par: Gong, Xiaochuan, et autres
Publié: (2025)
par: Gong, Xiaochuan, et autres
Publié: (2025)
Proactive DP: A Multple Target Optimization Framework for DP-SGD
par: van Dijk, Marten, et autres
Publié: (2021)
par: van Dijk, Marten, et autres
Publié: (2021)
Convergence Rate Analysis of the AdamW-Style Shampoo: Unifying One-Sided and Two-Sided Preconditioning
par: Li, Huan, et autres
Publié: (2026)
par: Li, Huan, et autres
Publié: (2026)
A Provably Convergent Plug-and-Play Framework for Stochastic Bilevel Optimization
par: Chu, Tianshu, et autres
Publié: (2025)
par: Chu, Tianshu, et autres
Publié: (2025)
Convergence of Clipped-SGD for Convex $(L_0,L_1)$-Smooth Optimization with Heavy-Tailed Noise
par: Chezhegov, Savelii, et autres
Publié: (2025)
par: Chezhegov, Savelii, et autres
Publié: (2025)
The Marginal Value of Momentum for Small Learning Rate SGD
par: Wang, Runzhe, et autres
Publié: (2023)
par: Wang, Runzhe, et autres
Publié: (2023)
Analyzing and Enhancing the Backward-Pass Convergence of Unrolled Optimization
par: Kotary, James, et autres
Publié: (2023)
par: Kotary, James, et autres
Publié: (2023)
A Theoretical and Empirical Study on the Convergence of Adam with an "Exact" Constant Step Size in Non-Convex Settings
par: Mazumder, Alokendu, et autres
Publié: (2023)
par: Mazumder, Alokendu, et autres
Publié: (2023)
On the $O(\frac{\sqrt{d}}{K^{1/4}})$ Convergence Rate of AdamW Measured by $\ell_1$ Norm
par: Li, Huan, et autres
Publié: (2025)
par: Li, Huan, et autres
Publié: (2025)
Convergence Guarantees for RMSProp and Adam in Generalized-smooth Non-convex Optimization with Affine Noise Variance
par: Zhang, Qi, et autres
Publié: (2024)
par: Zhang, Qi, et autres
Publié: (2024)
Asynchronous Decentralized SGD under Non-Convexity: A Block-Coordinate Descent Framework
par: Zhou, Yijie, et autres
Publié: (2025)
par: Zhou, Yijie, et autres
Publié: (2025)
Last-Iterate Convergence of Randomized Kaczmarz and SGD with Greedy Step Size
par: Dereziński, Michał, et autres
Publié: (2026)
par: Dereziński, Michał, et autres
Publié: (2026)
Documents similaires
-
The Rich and the Simple: On the Implicit Bias of Adam and SGD
par: Vasudeva, Bhavya, et autres
Publié: (2025) -
Adam-HNAG: A Convergent Reformulation of Adam with Accelerated Rate
par: Yu, Yaxin, et autres
Publié: (2026) -
Adam-SHANG: A Convergent Adam-Type Method for Stochastic Smooth Convex Optimization
par: Yu, Yaxin, et autres
Publié: (2026) -
Adam-family Methods for Nonsmooth Optimization with Convergence Guarantees
par: Xiao, Nachuan, et autres
Publié: (2023) -
Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling
par: Srećković, Teodora, et autres
Publié: (2025)