Convergence rates for the Adam optimizer
Fuente:
arXiv
Guardado en:
| Autores principales: | Dereich, Steffen, Jentzen, Arnulf |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
ODE approximation for the Adam algorithm: General and overparametrized setting
por: Dereich, Steffen, et al.
Publicado: (2025)
por: Dereich, Steffen, et al.
Publicado: (2025)
Averaged Adam accelerates stochastic optimization in the training of deep neural network approximations for partial differential equation and optimal control problems
por: Dereich, Steffen, et al.
Publicado: (2025)
por: Dereich, Steffen, et al.
Publicado: (2025)
Non-convergence of Adam and other adaptive stochastic gradient descent optimization methods for non-vanishing learning rates
por: Dereich, Steffen, et al.
Publicado: (2024)
por: Dereich, Steffen, et al.
Publicado: (2024)
Learning rate adaptive stochastic gradient descent optimization methods: numerical simulations for deep learning methods for partial differential equations and convergence analyses
por: Dereich, Steffen, et al.
Publicado: (2024)
por: Dereich, Steffen, et al.
Publicado: (2024)
Sharp higher order convergence rates for the Adam optimizer
por: Dereich, Steffen, et al.
Publicado: (2025)
por: Dereich, Steffen, et al.
Publicado: (2025)
On the existence of minimizers in shallow residual ReLU neural network optimization landscapes
por: Dereich, Steffen, et al.
Publicado: (2023)
por: Dereich, Steffen, et al.
Publicado: (2023)
Non-convergence to global minimizers for Adam and stochastic gradient descent optimization and constructions of local minimizers in the training of artificial neural networks
por: Jentzen, Arnulf, et al.
Publicado: (2024)
por: Jentzen, Arnulf, et al.
Publicado: (2024)
Uniform a priori bounds and error analysis for the Adam stochastic gradient descent optimization method
por: Dereich, Steffen, et al.
Publicado: (2026)
por: Dereich, Steffen, et al.
Publicado: (2026)
Adam symmetry theorem: characterization of the convergence of the stochastic Adam optimizer
por: Dereich, Steffen, et al.
Publicado: (2025)
por: Dereich, Steffen, et al.
Publicado: (2025)
PADAM: Parallel averaged Adam reduces the error for stochastic optimization in scientific machine learning
por: Jentzen, Arnulf, et al.
Publicado: (2025)
por: Jentzen, Arnulf, et al.
Publicado: (2025)
Asymptotic stability properties and a priori bounds for Adam and other gradient descent optimization methods
por: Dereich, Steffen, et al.
Publicado: (2025)
por: Dereich, Steffen, et al.
Publicado: (2025)
On bounds for norms of reparameterized ReLU artificial neural network parameters: sums of fractional powers of the Lipschitz norm control the network parameter vector
por: Jentzen, Arnulf, et al.
Publicado: (2022)
por: Jentzen, Arnulf, et al.
Publicado: (2022)
Convergence rate of Tsallis entropic regularized optimal transport
por: Suguro, Takeshi, et al.
Publicado: (2023)
por: Suguro, Takeshi, et al.
Publicado: (2023)
SAD Neural Networks: Divergent Gradient Flows and Asymptotic Optimality via o-minimal Structures
por: Kranz, Julian, et al.
Publicado: (2025)
por: Kranz, Julian, et al.
Publicado: (2025)
Convergence rates for gradient descent in the training of overparameterized artificial neural networks with piecewise affine activation
por: Jentzen, Arnulf, et al.
Publicado: (2021)
por: Jentzen, Arnulf, et al.
Publicado: (2021)
Non-convergence to global minimizers in data driven supervised deep learning: Adam and stochastic gradient descent optimization provably fail to converge to global minimizers in the training of deep neural networks with ReLU activation
por: Do, Thang, et al.
Publicado: (2024)
por: Do, Thang, et al.
Publicado: (2024)
A Generalization Result for Convergence in Learning-to-Optimize
por: Sucker, Michael, et al.
Publicado: (2024)
por: Sucker, Michael, et al.
Publicado: (2024)
Weak Convergence Analysis of Online Neural Actor-Critic Algorithms
por: Lam, Samuel Chun-Hei, et al.
Publicado: (2024)
por: Lam, Samuel Chun-Hei, et al.
Publicado: (2024)
Convergence of SGD for Training Neural Networks with Sliced Wasserstein Losses
por: Tanguy, Eloi
Publicado: (2023)
por: Tanguy, Eloi
Publicado: (2023)
Controlling the Flow: Stability and Convergence for Stochastic Gradient Descent with Decaying Regularization
por: Kassing, Sebastian, et al.
Publicado: (2025)
por: Kassing, Sebastian, et al.
Publicado: (2025)
Convergence of coordinate ascent variational inference for log-concave measures via optimal transport
por: Arnese, Manuel, et al.
Publicado: (2024)
por: Arnese, Manuel, et al.
Publicado: (2024)
Prelimit Coupling and Steady-State Convergence of Constant-stepsize Nonsmooth Contractive SA
por: Zhang, Yixuan, et al.
Publicado: (2024)
por: Zhang, Yixuan, et al.
Publicado: (2024)
Wasserstein Convergence of Score-based Generative Models under Semiconvexity and Discontinuous Gradients
por: Bruno, Stefano, et al.
Publicado: (2025)
por: Bruno, Stefano, et al.
Publicado: (2025)
Convergence of Actor-Critic Learning for Mean Field Games and Mean Field Control in Continuous Spaces
por: Fouque, Jean-Pierre, et al.
Publicado: (2025)
por: Fouque, Jean-Pierre, et al.
Publicado: (2025)
Convergence Error Analysis of Reflected Gradient Langevin Dynamics for Globally Optimizing Non-Convex Constrained Problems
por: Sato, Kanji, et al.
Publicado: (2022)
por: Sato, Kanji, et al.
Publicado: (2022)
Rates of Convergence in the Central Limit Theorem for Markov Chains, with an Application to TD Learning
por: Srikant, R.
Publicado: (2024)
por: Srikant, R.
Publicado: (2024)
Non-convex entropic mean-field optimization via Best Response flow
por: Lascu, Razvan-Andrei, et al.
Publicado: (2025)
por: Lascu, Razvan-Andrei, et al.
Publicado: (2025)
On propagation of chaos for the Fisher-Rao gradient flow in entropic mean-field optimization
por: Lazić, Petra, et al.
Publicado: (2026)
por: Lazić, Petra, et al.
Publicado: (2026)
Proximal optimal transport divergences
por: Baptista, Ricardo, et al.
Publicado: (2025)
por: Baptista, Ricardo, et al.
Publicado: (2025)
Convergence rate of random scan Coordinate Ascent Variational Inference under log-concavity
por: Lavenant, Hugo, et al.
Publicado: (2024)
por: Lavenant, Hugo, et al.
Publicado: (2024)
Adam-HNAG: A Convergent Reformulation of Adam with Accelerated Rate
por: Yu, Yaxin, et al.
Publicado: (2026)
por: Yu, Yaxin, et al.
Publicado: (2026)
On the existence of optimal shallow feedforward networks with ReLU activation
por: Dereich, Steffen, et al.
Publicado: (2023)
por: Dereich, Steffen, et al.
Publicado: (2023)
Convergence of linear programming hierarchies for Gibbs states of spin systems
por: Fawzi, Hamza, et al.
Publicado: (2025)
por: Fawzi, Hamza, et al.
Publicado: (2025)
Deep neural networks can provably solve Bellman equations for Markov decision processes without the curse of dimensionality
por: Jentzen, Arnulf, et al.
Publicado: (2025)
por: Jentzen, Arnulf, et al.
Publicado: (2025)
Adam-SHANG: A Convergent Adam-Type Method for Stochastic Smooth Convex Optimization
por: Yu, Yaxin, et al.
Publicado: (2026)
por: Yu, Yaxin, et al.
Publicado: (2026)
On the consistent reasoning paradox of intelligence and optimal trust in AI: The power of 'I don't know'
por: Bastounis, Alexander, et al.
Publicado: (2024)
por: Bastounis, Alexander, et al.
Publicado: (2024)
On Convergence of Adam for Stochastic Optimization under Relaxed Assumptions
por: Hong, Yusu, et al.
Publicado: (2024)
por: Hong, Yusu, et al.
Publicado: (2024)
Adam-family Methods for Nonsmooth Optimization with Convergence Guarantees
por: Xiao, Nachuan, et al.
Publicado: (2023)
por: Xiao, Nachuan, et al.
Publicado: (2023)
Adam Converges Without Any Modification On Update Rules
por: Zhang, Yushun, et al.
Publicado: (2026)
por: Zhang, Yushun, et al.
Publicado: (2026)
Decoupled Functional Central Limit Theorems for Two-Time-Scale Stochastic Approximation
por: Han, Yuze, et al.
Publicado: (2024)
por: Han, Yuze, et al.
Publicado: (2024)
Ejemplares similares
-
ODE approximation for the Adam algorithm: General and overparametrized setting
por: Dereich, Steffen, et al.
Publicado: (2025) -
Averaged Adam accelerates stochastic optimization in the training of deep neural network approximations for partial differential equation and optimal control problems
por: Dereich, Steffen, et al.
Publicado: (2025) -
Non-convergence of Adam and other adaptive stochastic gradient descent optimization methods for non-vanishing learning rates
por: Dereich, Steffen, et al.
Publicado: (2024) -
Learning rate adaptive stochastic gradient descent optimization methods: numerical simulations for deep learning methods for partial differential equations and convergence analyses
por: Dereich, Steffen, et al.
Publicado: (2024) -
Sharp higher order convergence rates for the Adam optimizer
por: Dereich, Steffen, et al.
Publicado: (2025)