UAdam: Unified Adam-Type Algorithmic Framework for Non-Convex Stochastic Optimization
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | , , , |
|---|---|
| Format: | Preprint |
| Publié: |
2023
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866917780607991808 |
|---|---|
| author | Jiang, Yiming Liu, Jinlan Xu, Dongpo Mandic, Danilo P. |
| author_facet | Jiang, Yiming Liu, Jinlan Xu, Dongpo Mandic, Danilo P. |
| contents | Adam-type algorithms have become a preferred choice for optimisation in the deep learning setting, however, despite success, their convergence is still not well understood. To this end, we introduce a unified framework for Adam-type algorithms (called UAdam). This is equipped with a general form of the second-order moment, which makes it possible to include Adam and its variants as special cases, such as NAdam, AMSGrad, AdaBound, AdaFom, and Adan. This is supported by a rigorous convergence analysis of UAdam in the non-convex stochastic setting, showing that UAdam converges to the neighborhood of stationary points with the rate of $\mathcal{O}(1/T)$. Furthermore, the size of neighborhood decreases as $β$ increases. Importantly, our analysis only requires the first-order momentum factor to be close enough to 1, without any restrictions on the second-order momentum factor. Theoretical results also show that vanilla Adam can converge by selecting appropriate hyperparameters, which provides a theoretical guarantee for the analysis, applications, and further developments of the whole class of Adam-type algorithms. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2305_05675 |
| institution | arXiv |
| publishDate | 2023 |
| record_format | arxiv |
| spellingShingle | UAdam: Unified Adam-Type Algorithmic Framework for Non-Convex Stochastic Optimization Jiang, Yiming Liu, Jinlan Xu, Dongpo Mandic, Danilo P. Machine Learning Numerical Analysis Optimization and Control Adam-type algorithms have become a preferred choice for optimisation in the deep learning setting, however, despite success, their convergence is still not well understood. To this end, we introduce a unified framework for Adam-type algorithms (called UAdam). This is equipped with a general form of the second-order moment, which makes it possible to include Adam and its variants as special cases, such as NAdam, AMSGrad, AdaBound, AdaFom, and Adan. This is supported by a rigorous convergence analysis of UAdam in the non-convex stochastic setting, showing that UAdam converges to the neighborhood of stationary points with the rate of $\mathcal{O}(1/T)$. Furthermore, the size of neighborhood decreases as $β$ increases. Importantly, our analysis only requires the first-order momentum factor to be close enough to 1, without any restrictions on the second-order momentum factor. Theoretical results also show that vanilla Adam can converge by selecting appropriate hyperparameters, which provides a theoretical guarantee for the analysis, applications, and further developments of the whole class of Adam-type algorithms. |
| title | UAdam: Unified Adam-Type Algorithmic Framework for Non-Convex Stochastic Optimization |
| topic | Machine Learning Numerical Analysis Optimization and Control |
| url | https://arxiv.org/abs/2305.05675 |