Non-convergence to global minimizers in data driven supervised deep learning: Adam and stochastic gradient descent optimization provably fail to converge to global minimizers in the training of deep neural networks with ReLU activation
Fuente:
arXiv
Saved in:
| Main Authors: | Do, Thang, Hannibal, Sonja, Jentzen, Arnulf |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Non-convergence to the optimal risk for Adam and stochastic gradient descent optimization in the training of deep neural networks
by: Do, Thang, et al.
Published: (2025)
by: Do, Thang, et al.
Published: (2025)
Non-convergence of Adam and other adaptive stochastic gradient descent optimization methods for non-vanishing learning rates
by: Dereich, Steffen, et al.
Published: (2024)
by: Dereich, Steffen, et al.
Published: (2024)
Uniform a priori bounds and error analysis for the Adam stochastic gradient descent optimization method
by: Dereich, Steffen, et al.
Published: (2026)
by: Dereich, Steffen, et al.
Published: (2026)
Adam symmetry theorem: characterization of the convergence of the stochastic Adam optimizer
by: Dereich, Steffen, et al.
Published: (2025)
by: Dereich, Steffen, et al.
Published: (2025)
Asymptotic stability properties and a priori bounds for Adam and other gradient descent optimization methods
by: Dereich, Steffen, et al.
Published: (2025)
by: Dereich, Steffen, et al.
Published: (2025)
On the speed of convergence of discrete Pickands constants to continuous ones
by: Bisewski, Krzysztof, et al.
Published: (2021)
by: Bisewski, Krzysztof, et al.
Published: (2021)
Fast reliable pricing and calibration of the rough Heston model
by: Boyarchenko, Svetlana, et al.
Published: (2025)
by: Boyarchenko, Svetlana, et al.
Published: (2025)
Deep neural networks with ReLU, leaky ReLU, and softplus activation provably overcome the curse of dimensionality for space-time solutions of semilinear partial differential equations
by: Ackermann, Julia, et al.
Published: (2024)
by: Ackermann, Julia, et al.
Published: (2024)
A characterisation of the continuum Gaussian free field in $d \geq 2$ dimensions
by: Aru, Juhan, et al.
Published: (2021)
by: Aru, Juhan, et al.
Published: (2021)
Strong convergence in the infinite horizon of numerical methods for stochastic delay differential equations
by: Wang, Yudong, et al.
Published: (2025)
by: Wang, Yudong, et al.
Published: (2025)
Weak convergence rates for temporal numerical approximations of stochastic wave equations with multiplicative noise
by: Cox, Sonja, et al.
Published: (2019)
by: Cox, Sonja, et al.
Published: (2019)
Efficient implicit-explicit sparse stochastic method for high dimensional semi-linear nonlocal diffusion equations
by: Sheng, Changtao, et al.
Published: (2025)
by: Sheng, Changtao, et al.
Published: (2025)
Correct implied volatility shapes and reliable pricing in the rough Heston model
by: Boyarchenko, Svetlana, et al.
Published: (2024)
by: Boyarchenko, Svetlana, et al.
Published: (2024)
Affine models with path-dependence under parameter uncertainty and their application in finance
by: Geuchen, Benedikt, et al.
Published: (2022)
by: Geuchen, Benedikt, et al.
Published: (2022)
A convergent scheme for the Bayesian filtering problem based on the Fokker--Planck equation and deep splitting
by: Bågmark, Kasper, et al.
Published: (2024)
by: Bågmark, Kasper, et al.
Published: (2024)
Deep neural networks with ReLU, leaky ReLU, and softplus activation provably overcome the curse of dimensionality for Kolmogorov partial differential equations with Lipschitz nonlinearities in the $L^p$-sense
by: Ackermann, Julia, et al.
Published: (2023)
by: Ackermann, Julia, et al.
Published: (2023)
Rough volatility, path-dependent PDEs and weak rates of convergence
by: Bonesini, Ofelia, et al.
Published: (2023)
by: Bonesini, Ofelia, et al.
Published: (2023)
Investigation and Development of the Methodologies for Simulating Self-similar Processes
by: Peng, Qidi, et al.
Published: (2025)
by: Peng, Qidi, et al.
Published: (2025)
Allowing for imprecision in the game-theoretic characterisation of the Poisson process
by: Erreygers, Alexander
Published: (2026)
by: Erreygers, Alexander
Published: (2026)
Uniform in time convergence of numerical schemes for stochastic differential equations via Strong Exponential stability: Euler methods, Split-Step and Tamed Schemes
by: Angeli, Letizia, et al.
Published: (2023)
by: Angeli, Letizia, et al.
Published: (2023)
Cover times with stochastic resetting
by: Linn, Samantha, et al.
Published: (2024)
by: Linn, Samantha, et al.
Published: (2024)
Convergence rate for random walk approximations of mean field BSDEs
by: Djehiche, Boualem, et al.
Published: (2024)
by: Djehiche, Boualem, et al.
Published: (2024)
The $C^{0,1}$ Itô-Ventzell formula for weak Dirichlet processes
by: Fießinger, Felix, et al.
Published: (2023)
by: Fießinger, Felix, et al.
Published: (2023)
Particle method for the numerical simulation of the path-dependent McKean-Vlasov equation
by: Bernou, Armand, et al.
Published: (2022)
by: Bernou, Armand, et al.
Published: (2022)
Multi-dimensional fractional Brownian motion in the G-setting
by: Biagini, Francesca, et al.
Published: (2023)
by: Biagini, Francesca, et al.
Published: (2023)
Extending the noise of splitting to its completion and stability of Brownian maxima
by: Vidmar, Matija, et al.
Published: (2024)
by: Vidmar, Matija, et al.
Published: (2024)
A deep implicit-explicit minimizing movement method for option pricing in jump-diffusion models
by: Georgoulis, Emmanuil H., et al.
Published: (2024)
by: Georgoulis, Emmanuil H., et al.
Published: (2024)
Stochastic Differential Equations Driven by G-Brownian Motion with Mean Reflections
by: Li, Hanwu, et al.
Published: (2023)
by: Li, Hanwu, et al.
Published: (2023)
Stochastic Domination of Exit Times for Random Walks and Brownian Motion with Drift
by: Geng, Xi, et al.
Published: (2024)
by: Geng, Xi, et al.
Published: (2024)
On explosion time in stochastic differential equations driven by fractional Brownian motion
by: Garzon, Johanna, et al.
Published: (2024)
by: Garzon, Johanna, et al.
Published: (2024)
A stochastic Galerkin method with adaptive time-stepping for the Navier-Stokes equations
by: Sousedík, Bedřich, et al.
Published: (2022)
by: Sousedík, Bedřich, et al.
Published: (2022)
Convergence to good non-optimal critical points in the training of neural networks: Gradient descent optimization with one random initialization overcomes all bad non-global local minima with high probability
by: Ibragimov, Shokhrukh, et al.
Published: (2022)
by: Ibragimov, Shokhrukh, et al.
Published: (2022)
Nested Optimal Transport Distances
by: Bontorno, Ruben, et al.
Published: (2025)
by: Bontorno, Ruben, et al.
Published: (2025)
Estimation of the Self-similarity Index of Non-stationary Increments Self-similar Processes via Lamperti Transformations
by: Wu, William, et al.
Published: (2026)
by: Wu, William, et al.
Published: (2026)
Dirichlet-Neumann Averaging: The DNA of Efficient Gaussian Process Simulation
by: Kutri, Robert, et al.
Published: (2024)
by: Kutri, Robert, et al.
Published: (2024)
Wasserstein error estimates between telegraph processes and Brownian motion
by: Barrera, Gerardo, et al.
Published: (2025)
by: Barrera, Gerardo, et al.
Published: (2025)
Data-Driven Stochastic Optimal Control for Intraday Electricity Trading by Renewable Producers
by: Hammouda, Chiheb Ben, et al.
Published: (2026)
by: Hammouda, Chiheb Ben, et al.
Published: (2026)
Sharp barrier estimates for Bessel bridges
by: Chiarini, Leandro, et al.
Published: (2025)
by: Chiarini, Leandro, et al.
Published: (2025)
Drawdowns of diffusions
by: Salminen, Paavo, et al.
Published: (2024)
by: Salminen, Paavo, et al.
Published: (2024)
Fully discrete approximation of the semilinear stochastic wave equation on the sphere
by: Cohen, David, et al.
Published: (2026)
by: Cohen, David, et al.
Published: (2026)
Similar Items
-
Non-convergence to the optimal risk for Adam and stochastic gradient descent optimization in the training of deep neural networks
by: Do, Thang, et al.
Published: (2025) -
Non-convergence of Adam and other adaptive stochastic gradient descent optimization methods for non-vanishing learning rates
by: Dereich, Steffen, et al.
Published: (2024) -
Uniform a priori bounds and error analysis for the Adam stochastic gradient descent optimization method
by: Dereich, Steffen, et al.
Published: (2026) -
Adam symmetry theorem: characterization of the convergence of the stochastic Adam optimizer
by: Dereich, Steffen, et al.
Published: (2025) -
Asymptotic stability properties and a priori bounds for Adam and other gradient descent optimization methods
by: Dereich, Steffen, et al.
Published: (2025)