Non-convergence to global minimizers for Adam and stochastic gradient descent optimization and constructions of local minimizers in the training of artificial neural networks
Fuente:
arXiv
Guardado en:
| Autores principales: | Jentzen, Arnulf, Riekert, Adrian |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Averaged Adam accelerates stochastic optimization in the training of deep neural network approximations for partial differential equation and optimal control problems
por: Dereich, Steffen, et al.
Publicado: (2025)
por: Dereich, Steffen, et al.
Publicado: (2025)
Learning rate adaptive stochastic gradient descent optimization methods: numerical simulations for deep learning methods for partial differential equations and convergence analyses
por: Dereich, Steffen, et al.
Publicado: (2024)
por: Dereich, Steffen, et al.
Publicado: (2024)
Non-convergence to the optimal risk for Adam and stochastic gradient descent optimization in the training of deep neural networks
por: Do, Thang, et al.
Publicado: (2025)
por: Do, Thang, et al.
Publicado: (2025)
Non-convergence to global minimizers in data driven supervised deep learning: Adam and stochastic gradient descent optimization provably fail to converge to global minimizers in the training of deep neural networks with ReLU activation
por: Do, Thang, et al.
Publicado: (2024)
por: Do, Thang, et al.
Publicado: (2024)
PADAM: Parallel averaged Adam reduces the error for stochastic optimization in scientific machine learning
por: Jentzen, Arnulf, et al.
Publicado: (2025)
por: Jentzen, Arnulf, et al.
Publicado: (2025)
Convergence to good non-optimal critical points in the training of neural networks: Gradient descent optimization with one random initialization overcomes all bad non-global local minima with high probability
por: Ibragimov, Shokhrukh, et al.
Publicado: (2022)
por: Ibragimov, Shokhrukh, et al.
Publicado: (2022)
Convergence rates for the Adam optimizer
por: Dereich, Steffen, et al.
Publicado: (2024)
por: Dereich, Steffen, et al.
Publicado: (2024)
Non-convergence of Adam and other adaptive stochastic gradient descent optimization methods for non-vanishing learning rates
por: Dereich, Steffen, et al.
Publicado: (2024)
por: Dereich, Steffen, et al.
Publicado: (2024)
On bounds for norms of reparameterized ReLU artificial neural network parameters: sums of fractional powers of the Lipschitz norm control the network parameter vector
por: Jentzen, Arnulf, et al.
Publicado: (2022)
por: Jentzen, Arnulf, et al.
Publicado: (2022)
Sharp higher order convergence rates for the Adam optimizer
por: Dereich, Steffen, et al.
Publicado: (2025)
por: Dereich, Steffen, et al.
Publicado: (2025)
Asymptotic stability properties and a priori bounds for Adam and other gradient descent optimization methods
por: Dereich, Steffen, et al.
Publicado: (2025)
por: Dereich, Steffen, et al.
Publicado: (2025)
On the existence of minimizers in shallow residual ReLU neural network optimization landscapes
por: Dereich, Steffen, et al.
Publicado: (2023)
por: Dereich, Steffen, et al.
Publicado: (2023)
Uniform a priori bounds and error analysis for the Adam stochastic gradient descent optimization method
por: Dereich, Steffen, et al.
Publicado: (2026)
por: Dereich, Steffen, et al.
Publicado: (2026)
ODE approximation for the Adam algorithm: General and overparametrized setting
por: Dereich, Steffen, et al.
Publicado: (2025)
por: Dereich, Steffen, et al.
Publicado: (2025)
Convergence rates for gradient descent in the training of overparameterized artificial neural networks with piecewise affine activation
por: Jentzen, Arnulf, et al.
Publicado: (2021)
por: Jentzen, Arnulf, et al.
Publicado: (2021)
Convergence of stochastic gradient descent under a local Lojasiewicz condition for deep neural networks
por: An, Jing, et al.
Publicado: (2023)
por: An, Jing, et al.
Publicado: (2023)
Adam symmetry theorem: characterization of the convergence of the stochastic Adam optimizer
por: Dereich, Steffen, et al.
Publicado: (2025)
por: Dereich, Steffen, et al.
Publicado: (2025)
Convergence of continuous-time stochastic gradient descent with applications to deep neural networks
por: Lugosi, Gabor, et al.
Publicado: (2024)
por: Lugosi, Gabor, et al.
Publicado: (2024)
Gradient descent provably escapes saddle points in the training of shallow ReLU networks
por: Cheridito, Patrick, et al.
Publicado: (2022)
por: Cheridito, Patrick, et al.
Publicado: (2022)
On the global convergence of gradient descent for wide shallow models with bounded nonlinearities
por: Petit, Romain, et al.
Publicado: (2026)
por: Petit, Romain, et al.
Publicado: (2026)
Non-asymptotic convergence analysis of the stochastic gradient Hamiltonian Monte Carlo algorithm with discontinuous stochastic gradient with applications to training of ReLU neural networks
por: Liang, Luxu, et al.
Publicado: (2024)
por: Liang, Luxu, et al.
Publicado: (2024)
The duality structure gradient descent algorithm: analysis and applications to neural networks
por: Flynn, Thomas
Publicado: (2017)
por: Flynn, Thomas
Publicado: (2017)
The late-stage training dynamics of (stochastic) subgradient descent on homogeneous neural networks
por: Schechtman, Sholom, et al.
Publicado: (2025)
por: Schechtman, Sholom, et al.
Publicado: (2025)
New logarithmic step size for stochastic gradient descent
por: Shamaee, M. Soheil, et al.
Publicado: (2024)
por: Shamaee, M. Soheil, et al.
Publicado: (2024)
Precise gradient descent training dynamics for finite-width multi-layer neural networks
por: Han, Qiyang, et al.
Publicado: (2025)
por: Han, Qiyang, et al.
Publicado: (2025)
Flattened one-bit stochastic gradient descent: compressed distributed optimization with controlled variance
por: Stollenwerk, Alexander, et al.
Publicado: (2024)
por: Stollenwerk, Alexander, et al.
Publicado: (2024)
Global convergence of gradient descent for phase retrieval
por: Fougereux, Théodore, et al.
Publicado: (2024)
por: Fougereux, Théodore, et al.
Publicado: (2024)
A stochastic gradient descent algorithm with random search directions
por: Gbaguidi, Eméric
Publicado: (2025)
por: Gbaguidi, Eméric
Publicado: (2025)
Almost sure convergence rates of stochastic gradient methods under gradient domination
por: Weissmann, Simon, et al.
Publicado: (2024)
por: Weissmann, Simon, et al.
Publicado: (2024)
Nonconvex optimization and convergence of stochastic gradient descent, and solution of asynchronous game
por: Buck, Kevin, et al.
Publicado: (2025)
por: Buck, Kevin, et al.
Publicado: (2025)
A stochastic gradient method for trilevel optimization
por: Giovannelli, Tommaso, et al.
Publicado: (2025)
por: Giovannelli, Tommaso, et al.
Publicado: (2025)
A short proof of near-linear convergence of adaptive gradient descent under fourth-order growth and convexity
por: Davis, Damek, et al.
Publicado: (2026)
por: Davis, Damek, et al.
Publicado: (2026)
Solving a class of stochastic optimal control problems by physics-informed neural networks
por: Jiao, Zhe, et al.
Publicado: (2024)
por: Jiao, Zhe, et al.
Publicado: (2024)
Dealing with unbounded gradients in stochastic saddle-point optimization
por: Neu, Gergely, et al.
Publicado: (2024)
por: Neu, Gergely, et al.
Publicado: (2024)
Gradient descent inference in empirical risk minimization
por: Han, Qiyang, et al.
Publicado: (2024)
por: Han, Qiyang, et al.
Publicado: (2024)
Fast Spawn\&Prune (FS\&P): Global convergence of stochastic conic particle gradient descent via birth/death process
por: De Castro, Yohann, et al.
Publicado: (2026)
por: De Castro, Yohann, et al.
Publicado: (2026)
SAD Neural Networks: Divergent Gradient Flows and Asymptotic Optimality via o-minimal Structures
por: Kranz, Julian, et al.
Publicado: (2025)
por: Kranz, Julian, et al.
Publicado: (2025)
Block majorization-minimization with diminishing radius for constrained nonsmooth nonconvex optimization
por: Lyu, Hanbaek, et al.
Publicado: (2020)
por: Lyu, Hanbaek, et al.
Publicado: (2020)
Convergence and complexity of block majorization-minimization for constrained block-Riemannian optimization
por: Li, Yuchen, et al.
Publicado: (2023)
por: Li, Yuchen, et al.
Publicado: (2023)
Convergence of gradient flow for learning convolutional neural networks
por: Diederen, Jona-Maria, et al.
Publicado: (2026)
por: Diederen, Jona-Maria, et al.
Publicado: (2026)
Ejemplares similares
-
Averaged Adam accelerates stochastic optimization in the training of deep neural network approximations for partial differential equation and optimal control problems
por: Dereich, Steffen, et al.
Publicado: (2025) -
Learning rate adaptive stochastic gradient descent optimization methods: numerical simulations for deep learning methods for partial differential equations and convergence analyses
por: Dereich, Steffen, et al.
Publicado: (2024) -
Non-convergence to the optimal risk for Adam and stochastic gradient descent optimization in the training of deep neural networks
por: Do, Thang, et al.
Publicado: (2025) -
Non-convergence to global minimizers in data driven supervised deep learning: Adam and stochastic gradient descent optimization provably fail to converge to global minimizers in the training of deep neural networks with ReLU activation
por: Do, Thang, et al.
Publicado: (2024) -
PADAM: Parallel averaged Adam reduces the error for stochastic optimization in scientific machine learning
por: Jentzen, Arnulf, et al.
Publicado: (2025)