Adam Converges Without Any Modification On Update Rules
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Yushun, Li, Bingran, Chen, Congliang, Luo, Zhi-Quan, Sun, Ruoyu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Provable Adaptivity of Adam under Non-uniform Smoothness
von: Wang, Bohan, et al.
Veröffentlicht: (2022)
von: Wang, Bohan, et al.
Veröffentlicht: (2022)
On the Convergence of Adam under Non-uniform Smoothness: Separability from SGDM and Beyond
von: Wang, Bohan, et al.
Veröffentlicht: (2024)
von: Wang, Bohan, et al.
Veröffentlicht: (2024)
A Single-Loop Smoothed Gradient Descent-Ascent Algorithm for Nonconvex-Concave Min-Max Problems
von: Zhang, Jiawei, et al.
Veröffentlicht: (2020)
von: Zhang, Jiawei, et al.
Veröffentlicht: (2020)
Towards Quantifying the Hessian Structure of Neural Networks
von: Dong, Zhaorui, et al.
Veröffentlicht: (2025)
von: Dong, Zhaorui, et al.
Veröffentlicht: (2025)
Adam-HNAG: A Convergent Reformulation of Adam with Accelerated Rate
von: Yu, Yaxin, et al.
Veröffentlicht: (2026)
von: Yu, Yaxin, et al.
Veröffentlicht: (2026)
Why Transformers Need Adam: A Hessian Perspective
von: Zhang, Yushun, et al.
Veröffentlicht: (2024)
von: Zhang, Yushun, et al.
Veröffentlicht: (2024)
Understanding Adam Optimizer via Online Learning of Updates: Adam is FTRL in Disguise
von: Ahn, Kwangjun, et al.
Veröffentlicht: (2024)
von: Ahn, Kwangjun, et al.
Veröffentlicht: (2024)
Adam-SHANG: A Convergent Adam-Type Method for Stochastic Smooth Convex Optimization
von: Yu, Yaxin, et al.
Veröffentlicht: (2026)
von: Yu, Yaxin, et al.
Veröffentlicht: (2026)
Convergence of Steepest Descent and Adam under Non-Uniform Smoothness
von: Vaswani, Sharan, et al.
Veröffentlicht: (2026)
von: Vaswani, Sharan, et al.
Veröffentlicht: (2026)
Convergence rates for the Adam optimizer
von: Dereich, Steffen, et al.
Veröffentlicht: (2024)
von: Dereich, Steffen, et al.
Veröffentlicht: (2024)
A Comprehensive Framework for Analyzing the Convergence of Adam: Bridging the Gap with SGD
von: Jin, Ruinan, et al.
Veröffentlicht: (2024)
von: Jin, Ruinan, et al.
Veröffentlicht: (2024)
Adam-family Methods for Nonsmooth Optimization with Convergence Guarantees
von: Xiao, Nachuan, et al.
Veröffentlicht: (2023)
von: Xiao, Nachuan, et al.
Veröffentlicht: (2023)
On Convergence of Adam for Stochastic Optimization under Relaxed Assumptions
von: Hong, Yusu, et al.
Veröffentlicht: (2024)
von: Hong, Yusu, et al.
Veröffentlicht: (2024)
How to Set $β_1, β_2$ in Adam: An Online Learning Perspective
von: Nguyen, Quan
Veröffentlicht: (2025)
von: Nguyen, Quan
Veröffentlicht: (2025)
HomeAdam: Adam and AdamW Algorithms Sometimes Go Home to Obtain Better Provable Generalization
von: Huang, Feihu, et al.
Veröffentlicht: (2026)
von: Huang, Feihu, et al.
Veröffentlicht: (2026)
On the Convergence of Adam-Type Algorithm for Bilevel Optimization under Unbounded Smoothness
von: Gong, Xiaochuan, et al.
Veröffentlicht: (2025)
von: Gong, Xiaochuan, et al.
Veröffentlicht: (2025)
Convergence for Discrete Parameter Update Schemes
von: Wilson, Paul, et al.
Veröffentlicht: (2025)
von: Wilson, Paul, et al.
Veröffentlicht: (2025)
Convergence Guarantees for RMSProp and Adam in Generalized-smooth Non-convex Optimization with Affine Noise Variance
von: Zhang, Qi, et al.
Veröffentlicht: (2024)
von: Zhang, Qi, et al.
Veröffentlicht: (2024)
Convergence Rate Analysis of the AdamW-Style Shampoo: Unifying One-Sided and Two-Sided Preconditioning
von: Li, Huan, et al.
Veröffentlicht: (2026)
von: Li, Huan, et al.
Veröffentlicht: (2026)
Convergence of Distributed Adaptive Optimization with Local Updates
von: Cheng, Ziheng, et al.
Veröffentlicht: (2024)
von: Cheng, Ziheng, et al.
Veröffentlicht: (2024)
Drop-Muon: Update Less, Converge Faster
von: Gruntkowska, Kaja, et al.
Veröffentlicht: (2025)
von: Gruntkowska, Kaja, et al.
Veröffentlicht: (2025)
On the $O(\frac{\sqrt{d}}{K^{1/4}})$ Convergence Rate of AdamW Measured by $\ell_1$ Norm
von: Li, Huan, et al.
Veröffentlicht: (2025)
von: Li, Huan, et al.
Veröffentlicht: (2025)
State estimations and noise identifications with intermittent corrupted observations via Bayesian variational inference
von: Sun, Peng, et al.
Veröffentlicht: (2026)
von: Sun, Peng, et al.
Veröffentlicht: (2026)
Efficient Penalty-Based Bilevel Methods: Improved Analysis, Novel Updates, and Flatness Condition
von: Jiang, Liuyuan, et al.
Veröffentlicht: (2025)
von: Jiang, Liuyuan, et al.
Veröffentlicht: (2025)
Dynamic Regret via Discounted-to-Dynamic Reduction with Applications to Curved Losses and Adam Optimizer
von: Xie, Yan-Feng, et al.
Veröffentlicht: (2026)
von: Xie, Yan-Feng, et al.
Veröffentlicht: (2026)
A Theoretical and Empirical Study on the Convergence of Adam with an "Exact" Constant Step Size in Non-Convex Settings
von: Mazumder, Alokendu, et al.
Veröffentlicht: (2023)
von: Mazumder, Alokendu, et al.
Veröffentlicht: (2023)
Adam on Local Time: Addressing Nonstationarity in RL with Relative Adam Timesteps
von: Ellis, Benjamin, et al.
Veröffentlicht: (2024)
von: Ellis, Benjamin, et al.
Veröffentlicht: (2024)
Continuous Q-Score Matching: Diffusion Guided Reinforcement Learning for Continuous-Time Control
von: Hua, Chengxiu, et al.
Veröffentlicht: (2025)
von: Hua, Chengxiu, et al.
Veröffentlicht: (2025)
Finite Horizon Optimization: Framework and Applications
von: Zhang, Yushun, et al.
Veröffentlicht: (2024)
von: Zhang, Yushun, et al.
Veröffentlicht: (2024)
Less is More: Convergence Benefits of Fewer Data Weight Updates over Longer Horizon
von: Das, Rudrajit, et al.
Veröffentlicht: (2026)
von: Das, Rudrajit, et al.
Veröffentlicht: (2026)
Projective Proximal Gradient Descent for A Class of Nonconvex Nonsmooth Optimization Problems: Fast Convergence Without Kurdyka-Lojasiewicz (KL) Property
von: Yang, Yingzhen, et al.
Veröffentlicht: (2023)
von: Yang, Yingzhen, et al.
Veröffentlicht: (2023)
Convergence of Adam for Non-convex Objectives: Relaxed Hyperparameters and Non-ergodic Case
von: He, Meixuan, et al.
Veröffentlicht: (2023)
von: He, Meixuan, et al.
Veröffentlicht: (2023)
The Effectiveness of Local Updates for Decentralized Learning under Data Heterogeneity
von: Wu, Tongle, et al.
Veröffentlicht: (2024)
von: Wu, Tongle, et al.
Veröffentlicht: (2024)
Feature Augmentation of GNNs for ILPs: Local Uniqueness Suffices
von: Han, Qingyu, et al.
Veröffentlicht: (2025)
von: Han, Qingyu, et al.
Veröffentlicht: (2025)
Incremental Gauss--Newton Methods with Superlinear Convergence Rates
von: Zhou, Zhiling, et al.
Veröffentlicht: (2024)
von: Zhou, Zhiling, et al.
Veröffentlicht: (2024)
Implicit Bias of AdamW: $\ell_\infty$ Norm Constrained Optimization
von: Xie, Shuo, et al.
Veröffentlicht: (2024)
von: Xie, Shuo, et al.
Veröffentlicht: (2024)
When GNNs meet symmetry in ILPs: an orbit-based feature augmentation approach
von: Chen, Qian, et al.
Veröffentlicht: (2025)
von: Chen, Qian, et al.
Veröffentlicht: (2025)
Optimistic Online-to-Batch Conversions for Accelerated Convergence and Universality
von: Yan, Yu-Hu, et al.
Veröffentlicht: (2025)
von: Yan, Yu-Hu, et al.
Veröffentlicht: (2025)
Sketch-and-Project Meets Newton Method: Global $\mathcal O(k^{-2})$ Convergence with Low-Rank Updates
von: Hanzely, Slavomír
Veröffentlicht: (2023)
von: Hanzely, Slavomír
Veröffentlicht: (2023)
Constrained Stochastic Spectral Preconditioning Converges for Nonconvex Objectives
von: Oikonomidis, Konstantinos, et al.
Veröffentlicht: (2026)
von: Oikonomidis, Konstantinos, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Provable Adaptivity of Adam under Non-uniform Smoothness
von: Wang, Bohan, et al.
Veröffentlicht: (2022) -
On the Convergence of Adam under Non-uniform Smoothness: Separability from SGDM and Beyond
von: Wang, Bohan, et al.
Veröffentlicht: (2024) -
A Single-Loop Smoothed Gradient Descent-Ascent Algorithm for Nonconvex-Concave Min-Max Problems
von: Zhang, Jiawei, et al.
Veröffentlicht: (2020) -
Towards Quantifying the Hessian Structure of Neural Networks
von: Dong, Zhaorui, et al.
Veröffentlicht: (2025) -
Adam-HNAG: A Convergent Reformulation of Adam with Accelerated Rate
von: Yu, Yaxin, et al.
Veröffentlicht: (2026)