Uniform a priori bounds and error analysis for the Adam stochastic gradient descent optimization method
Fuente:
arXiv
Saved in:
| Main Authors: | Dereich, Steffen, Do, Thang, Jentzen, Arnulf |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Adam symmetry theorem: characterization of the convergence of the stochastic Adam optimizer
by: Dereich, Steffen, et al.
Published: (2025)
by: Dereich, Steffen, et al.
Published: (2025)
Non-convergence to the optimal risk for Adam and stochastic gradient descent optimization in the training of deep neural networks
by: Do, Thang, et al.
Published: (2025)
by: Do, Thang, et al.
Published: (2025)
Asymptotic stability properties and a priori bounds for Adam and other gradient descent optimization methods
by: Dereich, Steffen, et al.
Published: (2025)
by: Dereich, Steffen, et al.
Published: (2025)
Sharp higher order convergence rates for the Adam optimizer
by: Dereich, Steffen, et al.
Published: (2025)
by: Dereich, Steffen, et al.
Published: (2025)
Non-convergence of Adam and other adaptive stochastic gradient descent optimization methods for non-vanishing learning rates
by: Dereich, Steffen, et al.
Published: (2024)
by: Dereich, Steffen, et al.
Published: (2024)
Non-convergence to global minimizers in data driven supervised deep learning: Adam and stochastic gradient descent optimization provably fail to converge to global minimizers in the training of deep neural networks with ReLU activation
by: Do, Thang, et al.
Published: (2024)
by: Do, Thang, et al.
Published: (2024)
An inexact infeasible arc-search interior-point method for linear optimization problems
by: Iida, Einosuke, et al.
Published: (2024)
by: Iida, Einosuke, et al.
Published: (2024)
An optimally fast objective-function-free minimization algorithm using random subspaces
by: Bellavia, S., et al.
Published: (2023)
by: Bellavia, S., et al.
Published: (2023)
Joint Pricing and Matching for Resource Allocation Platforms via Min-cost Flow Problem
by: Hikima, Yuya, et al.
Published: (2024)
by: Hikima, Yuya, et al.
Published: (2024)
Refining asymptotic complexity bounds for nonconvex optimization methods, including why steepest descent is $o(ε^{-2})$ rather than $\mathcal{O}(ε^{-2})$
by: Gratton, Serge, et al.
Published: (2024)
by: Gratton, Serge, et al.
Published: (2024)
Escaping Saddle Points via Curvature-Calibrated Perturbations: A Complete Analysis with Explicit Constants and Empirical Validation
by: Alpay, Faruk, et al.
Published: (2025)
by: Alpay, Faruk, et al.
Published: (2025)
Optimization Strategies for Variational Quantum Algorithms in Noisy Landscapes
by: Novák, Vojtěch, et al.
Published: (2025)
by: Novák, Vojtěch, et al.
Published: (2025)
An objective-function-free algorithm for nonconvex stochastic optimization with deterministic equality and inequality constraints
by: Gratton, S., et al.
Published: (2026)
by: Gratton, S., et al.
Published: (2026)
Learning the random variables in Monte Carlo simulations with stochastic gradient descent: Machine learning for parametric PDEs and financial derivative pricing
by: Becker, Sebastian, et al.
Published: (2022)
by: Becker, Sebastian, et al.
Published: (2022)
An infeasible interior-point arc-search method with Nesterov's restarting strategy for linear programming problems
by: Iida, Einosuke, et al.
Published: (2023)
by: Iida, Einosuke, et al.
Published: (2023)
A Riemannian gradient descent method for optimization on the indefinite Stiefel manifold
by: Van Tiep, Dinh, et al.
Published: (2024)
by: Van Tiep, Dinh, et al.
Published: (2024)
The non-intrusive reduced basis two-grid method applied to sensitivity analysis
by: Grosjean, Elise, et al.
Published: (2023)
by: Grosjean, Elise, et al.
Published: (2023)
Iteration complexity of the Difference-of-Convex Algorithm for unconstrained optimization: a simple proof
by: Gratton, Serge, et al.
Published: (2026)
by: Gratton, Serge, et al.
Published: (2026)
Examples of slow convergence for adaptive regularization optimization methods are not isolated
by: Toint, Philippe L.
Published: (2024)
by: Toint, Philippe L.
Published: (2024)
prunAdag: an adaptive pruning-aware gradient method
by: Porcelli, Margherita, et al.
Published: (2025)
by: Porcelli, Margherita, et al.
Published: (2025)
An objective-function-free algorithm for general smooth constrained optimization
by: Bellavia, S., et al.
Published: (2026)
by: Bellavia, S., et al.
Published: (2026)
Zeroth-order gradient estimators for stochastic problems with decision-dependent distributions
by: Hikima, Yuya, et al.
Published: (2025)
by: Hikima, Yuya, et al.
Published: (2025)
A Primal-Dual Frank-Wolfe Algorithm for Linear Programming
by: Hough, Matthew, et al.
Published: (2024)
by: Hough, Matthew, et al.
Published: (2024)
Sparse Training of Neural Networks based on Multilevel Mirror Descent
by: Lunk, Yannick, et al.
Published: (2026)
by: Lunk, Yannick, et al.
Published: (2026)
Nonlinear PageRank Problem for Local Graph Partitioning
by: Kodsi, Costy, et al.
Published: (2024)
by: Kodsi, Costy, et al.
Published: (2024)
A Stochastic Objective-Function-Free Adaptive Regularization Method with Optimal Complexity
by: Gratton, Serge, et al.
Published: (2024)
by: Gratton, Serge, et al.
Published: (2024)
Fast Stochastic Second-Order Adagrad for Nonconvex Bound-Constrained Optimization
by: Bellavia, S., et al.
Published: (2025)
by: Bellavia, S., et al.
Published: (2025)
Degrees-of-freedom penalized piecewise regression
by: Volz, Stefan, et al.
Published: (2023)
by: Volz, Stefan, et al.
Published: (2023)
Convergence to good non-optimal critical points in the training of neural networks: Gradient descent optimization with one random initialization overcomes all bad non-global local minima with high probability
by: Ibragimov, Shokhrukh, et al.
Published: (2022)
by: Ibragimov, Shokhrukh, et al.
Published: (2022)
A Simple First-Order Algorithm for Full-Rank Equality Constrained Optimization
by: Gratton, Serge, et al.
Published: (2025)
by: Gratton, Serge, et al.
Published: (2025)
Moduli space of optimization algorithms
by: Pasechnyuk-Vilensky, Dmitry, et al.
Published: (2025)
by: Pasechnyuk-Vilensky, Dmitry, et al.
Published: (2025)
NOVAK: Unified adaptive optimizer for deep neural networks
by: Kavun, Sergii
Published: (2026)
by: Kavun, Sergii
Published: (2026)
Handbook of Convergence Theorems for (Stochastic) Gradient Methods
by: Garrigos, Guillaume, et al.
Published: (2023)
by: Garrigos, Guillaume, et al.
Published: (2023)
Local properties of neural networks through the lens of layer-wise Hessians
by: Bolshim, Maxim, et al.
Published: (2025)
by: Bolshim, Maxim, et al.
Published: (2025)
Inter-Layer Hessian Analysis of Neural Networks with DAG Architectures
by: Bolshim, Maxim, et al.
Published: (2026)
by: Bolshim, Maxim, et al.
Published: (2026)
Stochastic Estimation of the Layer-wise Hessian Trace for Monitoring Neural-network Training
by: Bolshim, Maxim, et al.
Published: (2026)
by: Bolshim, Maxim, et al.
Published: (2026)
Gradient Descent Methods for Regularized Optimization
by: Nikolovski, Filip, et al.
Published: (2024)
by: Nikolovski, Filip, et al.
Published: (2024)
Majorization-Minimization-Based Levenberg--Marquardt Method for Constrained Nonlinear Least Squares
by: Marumo, Naoki, et al.
Published: (2020)
by: Marumo, Naoki, et al.
Published: (2020)
Where's Ben Nevis? A 2D optimisation benchmark with 957,174 local optima based on Great Britain terrain data
by: Wei, Yuhang, et al.
Published: (2024)
by: Wei, Yuhang, et al.
Published: (2024)
Data-Dependent Complexity of First-Order Methods for Binary Classification
by: Hough, Matthew, et al.
Published: (2025)
by: Hough, Matthew, et al.
Published: (2025)
Similar Items
-
Adam symmetry theorem: characterization of the convergence of the stochastic Adam optimizer
by: Dereich, Steffen, et al.
Published: (2025) -
Non-convergence to the optimal risk for Adam and stochastic gradient descent optimization in the training of deep neural networks
by: Do, Thang, et al.
Published: (2025) -
Asymptotic stability properties and a priori bounds for Adam and other gradient descent optimization methods
by: Dereich, Steffen, et al.
Published: (2025) -
Sharp higher order convergence rates for the Adam optimizer
by: Dereich, Steffen, et al.
Published: (2025) -
Non-convergence of Adam and other adaptive stochastic gradient descent optimization methods for non-vanishing learning rates
by: Dereich, Steffen, et al.
Published: (2024)