ODE approximation for the Adam algorithm: General and overparametrized setting
Fuente:
arXiv
Saved in:
| Main Authors: | Dereich, Steffen, Jentzen, Arnulf, Kassing, Sebastian |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Convergence rates for the Adam optimizer
by: Dereich, Steffen, et al.
Published: (2024)
by: Dereich, Steffen, et al.
Published: (2024)
Averaged Adam accelerates stochastic optimization in the training of deep neural network approximations for partial differential equation and optimal control problems
by: Dereich, Steffen, et al.
Published: (2025)
by: Dereich, Steffen, et al.
Published: (2025)
On the existence of minimizers in shallow residual ReLU neural network optimization landscapes
by: Dereich, Steffen, et al.
Published: (2023)
by: Dereich, Steffen, et al.
Published: (2023)
Learning rate adaptive stochastic gradient descent optimization methods: numerical simulations for deep learning methods for partial differential equations and convergence analyses
by: Dereich, Steffen, et al.
Published: (2024)
by: Dereich, Steffen, et al.
Published: (2024)
Non-convergence of Adam and other adaptive stochastic gradient descent optimization methods for non-vanishing learning rates
by: Dereich, Steffen, et al.
Published: (2024)
by: Dereich, Steffen, et al.
Published: (2024)
Exponential convergence rates for momentum stochastic gradient descent in the overparametrized setting
by: Gess, Benjamin, et al.
Published: (2023)
by: Gess, Benjamin, et al.
Published: (2023)
Non-convergence to global minimizers for Adam and stochastic gradient descent optimization and constructions of local minimizers in the training of artificial neural networks
by: Jentzen, Arnulf, et al.
Published: (2024)
by: Jentzen, Arnulf, et al.
Published: (2024)
Controlling the Flow: Stability and Convergence for Stochastic Gradient Descent with Decaying Regularization
by: Kassing, Sebastian, et al.
Published: (2025)
by: Kassing, Sebastian, et al.
Published: (2025)
On the existence of optimal shallow feedforward networks with ReLU activation
by: Dereich, Steffen, et al.
Published: (2023)
by: Dereich, Steffen, et al.
Published: (2023)
Uniform a priori bounds and error analysis for the Adam stochastic gradient descent optimization method
by: Dereich, Steffen, et al.
Published: (2026)
by: Dereich, Steffen, et al.
Published: (2026)
Sharp higher order convergence rates for the Adam optimizer
by: Dereich, Steffen, et al.
Published: (2025)
by: Dereich, Steffen, et al.
Published: (2025)
Adam symmetry theorem: characterization of the convergence of the stochastic Adam optimizer
by: Dereich, Steffen, et al.
Published: (2025)
by: Dereich, Steffen, et al.
Published: (2025)
PADAM: Parallel averaged Adam reduces the error for stochastic optimization in scientific machine learning
by: Jentzen, Arnulf, et al.
Published: (2025)
by: Jentzen, Arnulf, et al.
Published: (2025)
On bounds for norms of reparameterized ReLU artificial neural network parameters: sums of fractional powers of the Lipschitz norm control the network parameter vector
by: Jentzen, Arnulf, et al.
Published: (2022)
by: Jentzen, Arnulf, et al.
Published: (2022)
Asymptotic stability properties and a priori bounds for Adam and other gradient descent optimization methods
by: Dereich, Steffen, et al.
Published: (2025)
by: Dereich, Steffen, et al.
Published: (2025)
SAD Neural Networks: Divergent Gradient Flows and Asymptotic Optimality via o-minimal Structures
by: Kranz, Julian, et al.
Published: (2025)
by: Kranz, Julian, et al.
Published: (2025)
Convergence of stochastic gradient descent schemes for Lojasiewicz-landscapes
by: Dereich, Steffen, et al.
Published: (2021)
by: Dereich, Steffen, et al.
Published: (2021)
On the SAGA algorithm with decreasing step
by: Fredes, Luis, et al.
Published: (2024)
by: Fredes, Luis, et al.
Published: (2024)
Deep learning based numerical approximation algorithms for stochastic partial differential equations
by: Beck, Christian, et al.
Published: (2020)
by: Beck, Christian, et al.
Published: (2020)
A stochastic gradient descent algorithm with random search directions
by: Gbaguidi, Eméric
Published: (2025)
by: Gbaguidi, Eméric
Published: (2025)
Non-convergence to global minimizers in data driven supervised deep learning: Adam and stochastic gradient descent optimization provably fail to converge to global minimizers in the training of deep neural networks with ReLU activation
by: Do, Thang, et al.
Published: (2024)
by: Do, Thang, et al.
Published: (2024)
Hilbert's projective metric for functions of bounded growth and exponential convergence of Sinkhorn's algorithm
by: Eckstein, Stephan
Published: (2023)
by: Eckstein, Stephan
Published: (2023)
Polygonal Unadjusted Langevin Algorithms: Creating stable and efficient adaptive algorithms for neural networks
by: Lim, Dong-Young, et al.
Published: (2021)
by: Lim, Dong-Young, et al.
Published: (2021)
Function approximation by neural nets in the mean-field regime: Entropic regularization and controlled McKean-Vlasov dynamics
by: Tzen, Belinda, et al.
Published: (2020)
by: Tzen, Belinda, et al.
Published: (2020)
An abstract effective convergence theorem for stochastic processes, with applications to stochastic approximation
by: Neri, Morenikeji, et al.
Published: (2025)
by: Neri, Morenikeji, et al.
Published: (2025)
A Generalization Result for Convergence in Learning-to-Optimize
by: Sucker, Michael, et al.
Published: (2024)
by: Sucker, Michael, et al.
Published: (2024)
Generalized Wasserstein Flow Matching: Transport Plans, Everywhere, All at Once
by: Piening, Moritz, et al.
Published: (2026)
by: Piening, Moritz, et al.
Published: (2026)
Concentration of General Stochastic Approximation Under Heavy-Tailed Markovian Noise
by: Agrawal, Shubhada, et al.
Published: (2026)
by: Agrawal, Shubhada, et al.
Published: (2026)
Wasserstein Convergence of Score-based Generative Models under Semiconvexity and Discontinuous Gradients
by: Bruno, Stefano, et al.
Published: (2025)
by: Bruno, Stefano, et al.
Published: (2025)
Optimal Asymptotic Rates for (Stochastic) Gradient Descent under the Local PL-Condition: A Geometric Approach
by: Kassing, Sebastian, et al.
Published: (2026)
by: Kassing, Sebastian, et al.
Published: (2026)
Optimal and instance-dependent guarantees for Markovian linear stochastic approximation
by: Mou, Wenlong, et al.
Published: (2021)
by: Mou, Wenlong, et al.
Published: (2021)
On Bellman equations for continuous-time policy evaluation I: discretization and approximation
by: Mou, Wenlong, et al.
Published: (2024)
by: Mou, Wenlong, et al.
Published: (2024)
Deep neural networks can provably solve Bellman equations for Markov decision processes without the curse of dimensionality
by: Jentzen, Arnulf, et al.
Published: (2025)
by: Jentzen, Arnulf, et al.
Published: (2025)
Convergence rates for gradient descent in the training of overparameterized artificial neural networks with piecewise affine activation
by: Jentzen, Arnulf, et al.
Published: (2021)
by: Jentzen, Arnulf, et al.
Published: (2021)
Langevin dynamics based algorithm e-TH$\varepsilon$O POULA for stochastic optimization problems with discontinuous stochastic gradient
by: Lim, Dong-Young, et al.
Published: (2022)
by: Lim, Dong-Young, et al.
Published: (2022)
Space-time deep neural network approximations for high-dimensional partial differential equations
by: Hornung, Fabian, et al.
Published: (2020)
by: Hornung, Fabian, et al.
Published: (2020)
The Role of Target Update Frequencies in Q-Learning
by: Weissmann, Simon, et al.
Published: (2026)
by: Weissmann, Simon, et al.
Published: (2026)
Model Predictive Control is almost Optimal for Heterogeneous Restless Multi-armed Bandits
by: Narasimha, Dheeraj, et al.
Published: (2025)
by: Narasimha, Dheeraj, et al.
Published: (2025)
Non-convex entropic mean-field optimization via Best Response flow
by: Lascu, Razvan-Andrei, et al.
Published: (2025)
by: Lascu, Razvan-Andrei, et al.
Published: (2025)
Flatness-Aware Stochastic Gradient Langevin Dynamics
by: Bruno, Stefano, et al.
Published: (2025)
by: Bruno, Stefano, et al.
Published: (2025)
Similar Items
-
Convergence rates for the Adam optimizer
by: Dereich, Steffen, et al.
Published: (2024) -
Averaged Adam accelerates stochastic optimization in the training of deep neural network approximations for partial differential equation and optimal control problems
by: Dereich, Steffen, et al.
Published: (2025) -
On the existence of minimizers in shallow residual ReLU neural network optimization landscapes
by: Dereich, Steffen, et al.
Published: (2023) -
Learning rate adaptive stochastic gradient descent optimization methods: numerical simulations for deep learning methods for partial differential equations and convergence analyses
by: Dereich, Steffen, et al.
Published: (2024) -
Non-convergence of Adam and other adaptive stochastic gradient descent optimization methods for non-vanishing learning rates
by: Dereich, Steffen, et al.
Published: (2024)