Adam-SHANG: A Convergent Adam-Type Method for Stochastic Smooth Convex Optimization
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866918498620407808 |
|---|---|
| author | Yu, Yaxin Chen, Long Feng, Minfu |
| author_facet | Yu, Yaxin Chen, Long Feng, Minfu |
| contents | We propose Adam-SHANG, a Lyapunov-guided Adam-type method that couples momentum, adaptive preconditioning, and a curvature-aware correction through a more stable lagged-preconditioner update. For stochastic smooth convex optimization, we prove convergence in expectation under an admissible stepsize condition that can always be satisfied by a conservative spectral bound, without imposing global monotonicity on the second-moment sequence. To obtain a less conservative practical rule, we introduce a computable trace-ratio stepsize, motivated by a local coordinatewise alignment condition. The same structural update is also tested beyond the convex setting with simplified parameters. Experiments validate the predicted stochastic decay and show competitive training performance against Adam and AdamW on deep learning tasks. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2605_12878 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Adam-SHANG: A Convergent Adam-Type Method for Stochastic Smooth Convex Optimization Yu, Yaxin Chen, Long Feng, Minfu Optimization and Control Machine Learning We propose Adam-SHANG, a Lyapunov-guided Adam-type method that couples momentum, adaptive preconditioning, and a curvature-aware correction through a more stable lagged-preconditioner update. For stochastic smooth convex optimization, we prove convergence in expectation under an admissible stepsize condition that can always be satisfied by a conservative spectral bound, without imposing global monotonicity on the second-moment sequence. To obtain a less conservative practical rule, we introduce a computable trace-ratio stepsize, motivated by a local coordinatewise alignment condition. The same structural update is also tested beyond the convex setting with simplified parameters. Experiments validate the predicted stochastic decay and show competitive training performance against Adam and AdamW on deep learning tasks. |
| title | Adam-SHANG: A Convergent Adam-Type Method for Stochastic Smooth Convex Optimization |
| topic | Optimization and Control Machine Learning |
| url | https://arxiv.org/abs/2605.12878 |