Asymptotic stability properties and a priori bounds for Adam and other gradient descent optimization methods
Fuente:
arXiv
Saved in:
| Main Authors: | Dereich, Steffen, Graeber, Robin, Jentzen, Arnulf, Riekert, Adrian |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Non-convergence of Adam and other adaptive stochastic gradient descent optimization methods for non-vanishing learning rates
by: Dereich, Steffen, et al.
Published: (2024)
by: Dereich, Steffen, et al.
Published: (2024)
Uniform a priori bounds and error analysis for the Adam stochastic gradient descent optimization method
by: Dereich, Steffen, et al.
Published: (2026)
by: Dereich, Steffen, et al.
Published: (2026)
Non-convergence to the optimal risk for Adam and stochastic gradient descent optimization in the training of deep neural networks
by: Do, Thang, et al.
Published: (2025)
by: Do, Thang, et al.
Published: (2025)
Convergence to good non-optimal critical points in the training of neural networks: Gradient descent optimization with one random initialization overcomes all bad non-global local minima with high probability
by: Ibragimov, Shokhrukh, et al.
Published: (2022)
by: Ibragimov, Shokhrukh, et al.
Published: (2022)
Sharp higher order convergence rates for the Adam optimizer
by: Dereich, Steffen, et al.
Published: (2025)
by: Dereich, Steffen, et al.
Published: (2025)
Non-convergence to global minimizers in data driven supervised deep learning: Adam and stochastic gradient descent optimization provably fail to converge to global minimizers in the training of deep neural networks with ReLU activation
by: Do, Thang, et al.
Published: (2024)
by: Do, Thang, et al.
Published: (2024)
Adam symmetry theorem: characterization of the convergence of the stochastic Adam optimizer
by: Dereich, Steffen, et al.
Published: (2025)
by: Dereich, Steffen, et al.
Published: (2025)
Stochastic Mirror Descent for Convex Optimization with Consensus Constraints
by: Borovykh, Anastasia, et al.
Published: (2022)
by: Borovykh, Anastasia, et al.
Published: (2022)
Controlled fields, rough stochastic calculus, and Itô-Wentzell-Alekseev-Gröbner identities
by: Dause, Jannis R., et al.
Published: (2026)
by: Dause, Jannis R., et al.
Published: (2026)
Deep neural networks with ReLU, leaky ReLU, and softplus activation provably overcome the curse of dimensionality for space-time solutions of semilinear partial differential equations
by: Ackermann, Julia, et al.
Published: (2024)
by: Ackermann, Julia, et al.
Published: (2024)
Learning where to learn: Training data distribution optimization for scientific machine learning
by: Guerra, Nicolas, et al.
Published: (2025)
by: Guerra, Nicolas, et al.
Published: (2025)
Data-Free Asymptotics-Informed Operator Networks for Singularly Perturbed PDEs
by: Lee, Jinsil, et al.
Published: (2025)
by: Lee, Jinsil, et al.
Published: (2025)
Neural enrichment finite element method: A hybrid framework for problems with strong oscillations or interface problems
by: Guo, Shihan, et al.
Published: (2026)
by: Guo, Shihan, et al.
Published: (2026)
Deep neural networks with ReLU, leaky ReLU, and softplus activation provably overcome the curse of dimensionality for Kolmogorov partial differential equations with Lipschitz nonlinearities in the $L^p$-sense
by: Ackermann, Julia, et al.
Published: (2023)
by: Ackermann, Julia, et al.
Published: (2023)
Deep Predictor-Corrector Networks for Robust Parameter Estimation in Non-autonomous System with Discontinuous Inputs
by: Gu, Gyeongwan, et al.
Published: (2026)
by: Gu, Gyeongwan, et al.
Published: (2026)
Learning rate adaptive stochastic gradient descent optimization methods: numerical simulations for deep learning methods for partial differential equations and convergence analyses
by: Dereich, Steffen, et al.
Published: (2024)
by: Dereich, Steffen, et al.
Published: (2024)
Prediction of discretization of online GMsFEM using deep learning for Richards equation
by: Spiridonov, Denis, et al.
Published: (2024)
by: Spiridonov, Denis, et al.
Published: (2024)
An efficient gradient projection method for stochastic optimal control problem with expected integral state constraint
by: Wang, Qiming, et al.
Published: (2024)
by: Wang, Qiming, et al.
Published: (2024)
Neural Network Localized Orthogonal Decomposition for Numerical Homogenization of Diffusion Operators with Random Coefficients
by: Kröpfl, Fabian, et al.
Published: (2025)
by: Kröpfl, Fabian, et al.
Published: (2025)
Deep neural networks can provably solve Bellman equations for Markov decision processes without the curse of dimensionality
by: Jentzen, Arnulf, et al.
Published: (2025)
by: Jentzen, Arnulf, et al.
Published: (2025)
Divergence-Kernel method for scores of random systems
by: Ni, Angxiu
Published: (2025)
by: Ni, Angxiu
Published: (2025)
Online minimum search for a Brownian bridge
by: Wu, Erik, et al.
Published: (2024)
by: Wu, Erik, et al.
Published: (2024)
Spatio-temporal probabilistic forecast using MMAF-guided learning
by: Bardi, Leonardo, et al.
Published: (2026)
by: Bardi, Leonardo, et al.
Published: (2026)
A neural network method for scalar conservation laws with convergence rates for shock-wave solutions
by: Cao, Jiachuan, et al.
Published: (2026)
by: Cao, Jiachuan, et al.
Published: (2026)
Convergence of stochastic gradient descent schemes for Lojasiewicz-landscapes
by: Dereich, Steffen, et al.
Published: (2021)
by: Dereich, Steffen, et al.
Published: (2021)
On the existence of minimizers in shallow residual ReLU neural network optimization landscapes
by: Dereich, Steffen, et al.
Published: (2023)
by: Dereich, Steffen, et al.
Published: (2023)
Deep Operator BSDE: a Numerical Scheme to Approximate Solution Operators
by: Lozano, Pere Díaz, et al.
Published: (2024)
by: Lozano, Pere Díaz, et al.
Published: (2024)
A network based approach for unbalanced optimal transport on surfaces
by: Pan, Jiangong, et al.
Published: (2024)
by: Pan, Jiangong, et al.
Published: (2024)
Cover times of many diffusive or subdiffusive searchers
by: Kim, Hyunjoong, et al.
Published: (2023)
by: Kim, Hyunjoong, et al.
Published: (2023)
Weak Adversarial Neural Pushforward Method for the McKean-Vlasov / Mean-Field Fokker-Planck Equation
by: He, Andrew Qing, et al.
Published: (2026)
by: He, Andrew Qing, et al.
Published: (2026)
On finding optimal collective variables for complex systems by minimizing the deviation between effective and full dynamics
by: Zhang, Wei, et al.
Published: (2024)
by: Zhang, Wei, et al.
Published: (2024)
Fractional-Boundary-Regularized Deep Galerkin Method for Variational Inequalities in Mixed Optimal Stopping and Control
by: Zhao, Yun, et al.
Published: (2025)
by: Zhao, Yun, et al.
Published: (2025)
Random neural networks for rough volatility
by: Jacquier, Antoine, et al.
Published: (2023)
by: Jacquier, Antoine, et al.
Published: (2023)
Entropic Semi-Martingale Optimal Transport
by: Benamou, Jean-David, et al.
Published: (2024)
by: Benamou, Jean-David, et al.
Published: (2024)
Convergence of gradient descent for deep neural networks
by: Chatterjee, Sourav
Published: (2022)
by: Chatterjee, Sourav
Published: (2022)
Guaranteed lower bounds for cost functionals of time-periodic parabolic optimization problems
by: Wolfmayr, Monika
Published: (2019)
by: Wolfmayr, Monika
Published: (2019)
In almost all shallow analytic neural network optimization landscapes, efficient minimizers have strongly convex neighborhoods
by: Benning, Felix, et al.
Published: (2025)
by: Benning, Felix, et al.
Published: (2025)
Averaged Adam accelerates stochastic optimization in the training of deep neural network approximations for partial differential equation and optimal control problems
by: Dereich, Steffen, et al.
Published: (2025)
by: Dereich, Steffen, et al.
Published: (2025)
Consensus-based algorithms for stochastic optimization problems
by: Bonandin, Sabrina, et al.
Published: (2024)
by: Bonandin, Sabrina, et al.
Published: (2024)
Stochastic Galerkin methods for linear stability analysis of systems with parametric uncertainty
by: Sousedík, Bedřich, et al.
Published: (2022)
by: Sousedík, Bedřich, et al.
Published: (2022)
Similar Items
-
Non-convergence of Adam and other adaptive stochastic gradient descent optimization methods for non-vanishing learning rates
by: Dereich, Steffen, et al.
Published: (2024) -
Uniform a priori bounds and error analysis for the Adam stochastic gradient descent optimization method
by: Dereich, Steffen, et al.
Published: (2026) -
Non-convergence to the optimal risk for Adam and stochastic gradient descent optimization in the training of deep neural networks
by: Do, Thang, et al.
Published: (2025) -
Convergence to good non-optimal critical points in the training of neural networks: Gradient descent optimization with one random initialization overcomes all bad non-global local minima with high probability
by: Ibragimov, Shokhrukh, et al.
Published: (2022) -
Sharp higher order convergence rates for the Adam optimizer
by: Dereich, Steffen, et al.
Published: (2025)