Sharp higher order convergence rates for the Adam optimizer
Fuente:
arXiv
Saved in:
| Main Authors: | Dereich, Steffen, Jentzen, Arnulf, Riekert, Adrian |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Adam symmetry theorem: characterization of the convergence of the stochastic Adam optimizer
by: Dereich, Steffen, et al.
Published: (2025)
by: Dereich, Steffen, et al.
Published: (2025)
Uniform a priori bounds and error analysis for the Adam stochastic gradient descent optimization method
by: Dereich, Steffen, et al.
Published: (2026)
by: Dereich, Steffen, et al.
Published: (2026)
Convergence to good non-optimal critical points in the training of neural networks: Gradient descent optimization with one random initialization overcomes all bad non-global local minima with high probability
by: Ibragimov, Shokhrukh, et al.
Published: (2022)
by: Ibragimov, Shokhrukh, et al.
Published: (2022)
Asymptotic stability properties and a priori bounds for Adam and other gradient descent optimization methods
by: Dereich, Steffen, et al.
Published: (2025)
by: Dereich, Steffen, et al.
Published: (2025)
Non-convergence of Adam and other adaptive stochastic gradient descent optimization methods for non-vanishing learning rates
by: Dereich, Steffen, et al.
Published: (2024)
by: Dereich, Steffen, et al.
Published: (2024)
Preconditioned primal-dual dynamics in convex optimization: non-ergodic convergence rates
by: Apidopoulos, Vassilis, et al.
Published: (2025)
by: Apidopoulos, Vassilis, et al.
Published: (2025)
Fast Reflected Forward-Backward algorithm: achieving fast convergence rates for convex optimization with linear cone constraints
by: Bot, Radu Ioan, et al.
Published: (2024)
by: Bot, Radu Ioan, et al.
Published: (2024)
A Random Active Set Method for Strictly Convex Quadratic Problem with Simple Bounds
by: Gu, Ran, et al.
Published: (2021)
by: Gu, Ran, et al.
Published: (2021)
An adaptive framework for first-order gradient methods
by: Hu, Xiaozhe, et al.
Published: (2026)
by: Hu, Xiaozhe, et al.
Published: (2026)
Convergence of iterates and improved rates for accelerated augmented Lagrangian methods for linearly constrained convex optimization
by: He, Xin, et al.
Published: (2026)
by: He, Xin, et al.
Published: (2026)
Essential Convergence Rates of Continuous-Time Models for Optimization Methods
by: Ushiyama, Kansei, et al.
Published: (2025)
by: Ushiyama, Kansei, et al.
Published: (2025)
Bregman proximal gradient method for linear optimization under entropic constraints
by: Briceño-Arias, Luis M., et al.
Published: (2025)
by: Briceño-Arias, Luis M., et al.
Published: (2025)
Optimization in Theory and Practice
by: Wright, Stephen J.
Published: (2025)
by: Wright, Stephen J.
Published: (2025)
On the convergence of proximal gradient methods for convex simple bilevel optimization
by: Latafat, Puya, et al.
Published: (2023)
by: Latafat, Puya, et al.
Published: (2023)
Monomial barrier functions for the box-constrained convex optimization problems
by: Fayed, Hatem
Published: (2024)
by: Fayed, Hatem
Published: (2024)
Distributed Gradient-Regularized Newton Method: Scheduled Consensus and O(epsilon^{-1}) Global Iteration Complexity
by: Hu, Wei, et al.
Published: (2026)
by: Hu, Wei, et al.
Published: (2026)
Deep neural networks can provably solve Bellman equations for Markov decision processes without the curse of dimensionality
by: Jentzen, Arnulf, et al.
Published: (2025)
by: Jentzen, Arnulf, et al.
Published: (2025)
Dynamic FISTA for Convex Composite Bi-Level Optimization
by: Merchav, Roey, et al.
Published: (2024)
by: Merchav, Roey, et al.
Published: (2024)
Accelerated primal dual fixed point algorithm
by: Zhu, Ya-Nan
Published: (2025)
by: Zhu, Ya-Nan
Published: (2025)
A Proximal-Gradient Method for Solving Regularized Optimization Problems with General Constraints
by: Curtis, Frank E., et al.
Published: (2025)
by: Curtis, Frank E., et al.
Published: (2025)
A Proximal-Gradient Method for Constrained Optimization
by: Dai, Yutong, et al.
Published: (2024)
by: Dai, Yutong, et al.
Published: (2024)
On Solution Uniqueness and Robust Recovery for Sparse Regularization with a Gauge: from Dual Point of View
by: He, Jiahuan, et al.
Published: (2023)
by: He, Jiahuan, et al.
Published: (2023)
Preconditioned subgradient method for composite optimization: overparameterization and fast convergence
by: Díaz, Mateo, et al.
Published: (2025)
by: Díaz, Mateo, et al.
Published: (2025)
Minimization Over the Nonconvex Sparsity Constraint Using A Hybrid First-order method
by: Yang, Xiangyu, et al.
Published: (2021)
by: Yang, Xiangyu, et al.
Published: (2021)
Riemannian Adaptive Regularized Newton Methods with Hölder Continuous Hessians
by: Zhang, Chenyu, et al.
Published: (2023)
by: Zhang, Chenyu, et al.
Published: (2023)
A practical randomized trust-region method to escape saddle points in high dimension
by: Dragomir, Radu-Alexandru, et al.
Published: (2026)
by: Dragomir, Radu-Alexandru, et al.
Published: (2026)
Preconditioned Proximal Gradient Methods with Conjugate Momentum: A Subspace Perspective
by: Chen, Jian, et al.
Published: (2026)
by: Chen, Jian, et al.
Published: (2026)
Lipschitz continuity of solution multifunctions of extended $\ell_1$ regularization problems
by: Meng, Kaiwen, et al.
Published: (2024)
by: Meng, Kaiwen, et al.
Published: (2024)
Simplex Frank-Wolfe: Linear Convergence and Its Numerical Efficiency for Convex Optimization over Polytopes
by: Wang, Haoning, et al.
Published: (2025)
by: Wang, Haoning, et al.
Published: (2025)
Performance Estimation for Smooth and Strongly Convex Sets
by: Luner, Alan, et al.
Published: (2024)
by: Luner, Alan, et al.
Published: (2024)
A Regression-Based Prediction-Correction Method for Stochastic Time-Varying Optimization Problems
by: Kamijima, Tomoya, et al.
Published: (2025)
by: Kamijima, Tomoya, et al.
Published: (2025)
Stochastic Block Bregman Projection with Polyak-like Stepsize for Possibly Inconsistent Convex Feasibility Problems
by: Zhang, Lu, et al.
Published: (2026)
by: Zhang, Lu, et al.
Published: (2026)
On Averaging and Extrapolation for Gradient Descent
by: Luner, Alan, et al.
Published: (2024)
by: Luner, Alan, et al.
Published: (2024)
Extragradient method with feasible inexact projection to variational inequality problem
by: Millán, R. Díaz, et al.
Published: (2023)
by: Millán, R. Díaz, et al.
Published: (2023)
Tikhonov regularization of monotone operator flows not only ensures strong convergence of the trajectories but also speeds up the vanishing of the residuals
by: Bot, Radu Ioan, et al.
Published: (2024)
by: Bot, Radu Ioan, et al.
Published: (2024)
Relaxed and inertial nonlinear Forward-Backward algorithm
by: Maulén, Juan José, et al.
Published: (2025)
by: Maulén, Juan José, et al.
Published: (2025)
A Three-Operator Splitting Scheme Derived from Three-Block ADMM
by: Anshika, Anshika, et al.
Published: (2024)
by: Anshika, Anshika, et al.
Published: (2024)
Relaxed and Inertial Nonlinear Forward-Backward with Momentum
by: Roldán, Fernando, et al.
Published: (2024)
by: Roldán, Fernando, et al.
Published: (2024)
Forward Primal-Dual Half-Forward Algorithm for Splitting Four Operators
by: Roldán, Fernando
Published: (2023)
by: Roldán, Fernando
Published: (2023)
Technical results on the convergence of quasi-Newton methods for nonsmooth optimization
by: Gebken, Bennet
Published: (2025)
by: Gebken, Bennet
Published: (2025)
Similar Items
-
Adam symmetry theorem: characterization of the convergence of the stochastic Adam optimizer
by: Dereich, Steffen, et al.
Published: (2025) -
Uniform a priori bounds and error analysis for the Adam stochastic gradient descent optimization method
by: Dereich, Steffen, et al.
Published: (2026) -
Convergence to good non-optimal critical points in the training of neural networks: Gradient descent optimization with one random initialization overcomes all bad non-global local minima with high probability
by: Ibragimov, Shokhrukh, et al.
Published: (2022) -
Asymptotic stability properties and a priori bounds for Adam and other gradient descent optimization methods
by: Dereich, Steffen, et al.
Published: (2025) -
Non-convergence of Adam and other adaptive stochastic gradient descent optimization methods for non-vanishing learning rates
by: Dereich, Steffen, et al.
Published: (2024)