Learning rate adaptive stochastic gradient descent optimization methods: numerical simulations for deep learning methods for partial differential equations and convergence analyses
Fuente:
arXiv
Saved in:
| Main Authors: | Dereich, Steffen, Jentzen, Arnulf, Riekert, Adrian |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Averaged Adam accelerates stochastic optimization in the training of deep neural network approximations for partial differential equation and optimal control problems
by: Dereich, Steffen, et al.
Published: (2025)
by: Dereich, Steffen, et al.
Published: (2025)
Non-convergence of Adam and other adaptive stochastic gradient descent optimization methods for non-vanishing learning rates
by: Dereich, Steffen, et al.
Published: (2024)
by: Dereich, Steffen, et al.
Published: (2024)
Asymptotic stability properties and a priori bounds for Adam and other gradient descent optimization methods
by: Dereich, Steffen, et al.
Published: (2025)
by: Dereich, Steffen, et al.
Published: (2025)
Non-convergence to global minimizers for Adam and stochastic gradient descent optimization and constructions of local minimizers in the training of artificial neural networks
by: Jentzen, Arnulf, et al.
Published: (2024)
by: Jentzen, Arnulf, et al.
Published: (2024)
Sharp higher order convergence rates for the Adam optimizer
by: Dereich, Steffen, et al.
Published: (2025)
by: Dereich, Steffen, et al.
Published: (2025)
PADAM: Parallel averaged Adam reduces the error for stochastic optimization in scientific machine learning
by: Jentzen, Arnulf, et al.
Published: (2025)
by: Jentzen, Arnulf, et al.
Published: (2025)
Convergence rates for the Adam optimizer
by: Dereich, Steffen, et al.
Published: (2024)
by: Dereich, Steffen, et al.
Published: (2024)
Uniform a priori bounds and error analysis for the Adam stochastic gradient descent optimization method
by: Dereich, Steffen, et al.
Published: (2026)
by: Dereich, Steffen, et al.
Published: (2026)
Non-convergence to the optimal risk for Adam and stochastic gradient descent optimization in the training of deep neural networks
by: Do, Thang, et al.
Published: (2025)
by: Do, Thang, et al.
Published: (2025)
On the existence of minimizers in shallow residual ReLU neural network optimization landscapes
by: Dereich, Steffen, et al.
Published: (2023)
by: Dereich, Steffen, et al.
Published: (2023)
Non-convergence to global minimizers in data driven supervised deep learning: Adam and stochastic gradient descent optimization provably fail to converge to global minimizers in the training of deep neural networks with ReLU activation
by: Do, Thang, et al.
Published: (2024)
by: Do, Thang, et al.
Published: (2024)
Algorithmically Designed Artificial Neural Networks (ADANNs): Higher order deep operator learning for parametric partial differential equations
by: Jentzen, Arnulf, et al.
Published: (2023)
by: Jentzen, Arnulf, et al.
Published: (2023)
Martingale deep learning for very high dimensional quasi-linear partial differential equations and stochastic optimal controls
by: Cai, Wei, et al.
Published: (2024)
by: Cai, Wei, et al.
Published: (2024)
ODE approximation for the Adam algorithm: General and overparametrized setting
by: Dereich, Steffen, et al.
Published: (2025)
by: Dereich, Steffen, et al.
Published: (2025)
Mathematical analysis of the gradients in deep learning
by: Dereich, Steffen, et al.
Published: (2025)
by: Dereich, Steffen, et al.
Published: (2025)
Convergence to good non-optimal critical points in the training of neural networks: Gradient descent optimization with one random initialization overcomes all bad non-global local minima with high probability
by: Ibragimov, Shokhrukh, et al.
Published: (2022)
by: Ibragimov, Shokhrukh, et al.
Published: (2022)
Adam symmetry theorem: characterization of the convergence of the stochastic Adam optimizer
by: Dereich, Steffen, et al.
Published: (2025)
by: Dereich, Steffen, et al.
Published: (2025)
An accelerated gradient method with adaptive restart for convex multiobjective optimization problems
by: Luo, Hao, et al.
Published: (2025)
by: Luo, Hao, et al.
Published: (2025)
A perturbed preconditioned gradient descent method for the unconstrained minimization of composite objectives
by: Park, Jea-Hyun, et al.
Published: (2025)
by: Park, Jea-Hyun, et al.
Published: (2025)
Gauss-Southwell type descent methods for low-rank matrix optimization
by: Olikier, Guillaume, et al.
Published: (2023)
by: Olikier, Guillaume, et al.
Published: (2023)
Deep learning based numerical approximation algorithms for stochastic partial differential equations
by: Beck, Christian, et al.
Published: (2020)
by: Beck, Christian, et al.
Published: (2020)
Limited memory gradient methods for unconstrained optimization
by: Ferrandi, Giulia, et al.
Published: (2023)
by: Ferrandi, Giulia, et al.
Published: (2023)
Flattened one-bit stochastic gradient descent: compressed distributed optimization with controlled variance
by: Stollenwerk, Alexander, et al.
Published: (2024)
by: Stollenwerk, Alexander, et al.
Published: (2024)
A monotone block coordinate descent method for solving absolute value equations
by: Luo, Tingting, et al.
Published: (2024)
by: Luo, Tingting, et al.
Published: (2024)
On the convergence of stochastic variance reduced gradient for linear inverse problems
by: Jin, Bangti, et al.
Published: (2025)
by: Jin, Bangti, et al.
Published: (2025)
Error analysis for stochastic gradient optimization schemes using modified equations
by: Bréhier, Charles-Edouard, et al.
Published: (2024)
by: Bréhier, Charles-Edouard, et al.
Published: (2024)
Iterative solvers for partial differential equations with dissipative structure: Operator preconditioning and optimal control
by: Mehrmann, Volker, et al.
Published: (2025)
by: Mehrmann, Volker, et al.
Published: (2025)
Stochastic dual coordinate descent with adaptive heavy ball momentum for linearly constrained convex optimization
by: Zeng, Yun, et al.
Published: (2023)
by: Zeng, Yun, et al.
Published: (2023)
On the convergence analysis of the decentralized projected gradient descent method
by: Choi, Woocheol, et al.
Published: (2023)
by: Choi, Woocheol, et al.
Published: (2023)
Weak convergence rates for temporal numerical approximations of stochastic wave equations with multiplicative noise
by: Cox, Sonja, et al.
Published: (2019)
by: Cox, Sonja, et al.
Published: (2019)
Nonconvex optimization and convergence of stochastic gradient descent, and solution of asynchronous game
by: Buck, Kevin, et al.
Published: (2025)
by: Buck, Kevin, et al.
Published: (2025)
On the existence of optimal shallow feedforward networks with ReLU activation
by: Dereich, Steffen, et al.
Published: (2023)
by: Dereich, Steffen, et al.
Published: (2023)
Optimal local linear convergence of Nesterov's accelerated gradient method for $C^2$ functions under the Polyak--Łojasiewicz inequality
by: Feng, Zixu, et al.
Published: (2026)
by: Feng, Zixu, et al.
Published: (2026)
Convergence rates for gradient descent in the training of overparameterized artificial neural networks with piecewise affine activation
by: Jentzen, Arnulf, et al.
Published: (2021)
by: Jentzen, Arnulf, et al.
Published: (2021)
An unfitted finite element method for PDE-constrained shape optimization via shape gradient flow
by: Gong, Wei, et al.
Published: (2026)
by: Gong, Wei, et al.
Published: (2026)
Almost sure convergence of stochastic Hamiltonian descent methods
by: Williamson, Måns, et al.
Published: (2024)
by: Williamson, Måns, et al.
Published: (2024)
Convergence of the deep BSDE method for stochastic control problems formulated through the stochastic maximum principle
by: Huang, Zhipeng, et al.
Published: (2024)
by: Huang, Zhipeng, et al.
Published: (2024)
A Primal-dual hybrid gradient method for solving optimal control problems and the corresponding Hamilton-Jacobi PDEs
by: Meng, Tingwei, et al.
Published: (2024)
by: Meng, Tingwei, et al.
Published: (2024)
Neural incomplete factorization: learning preconditioners for the conjugate gradient method
by: Häusner, Paul, et al.
Published: (2023)
by: Häusner, Paul, et al.
Published: (2023)
An augmented Lagrangian trust-region method with inexact gradient evaluations to accelerate constrained optimization problems using model hyperreduction
by: Wen, Tianshu, et al.
Published: (2024)
by: Wen, Tianshu, et al.
Published: (2024)
Similar Items
-
Averaged Adam accelerates stochastic optimization in the training of deep neural network approximations for partial differential equation and optimal control problems
by: Dereich, Steffen, et al.
Published: (2025) -
Non-convergence of Adam and other adaptive stochastic gradient descent optimization methods for non-vanishing learning rates
by: Dereich, Steffen, et al.
Published: (2024) -
Asymptotic stability properties and a priori bounds for Adam and other gradient descent optimization methods
by: Dereich, Steffen, et al.
Published: (2025) -
Non-convergence to global minimizers for Adam and stochastic gradient descent optimization and constructions of local minimizers in the training of artificial neural networks
by: Jentzen, Arnulf, et al.
Published: (2024) -
Sharp higher order convergence rates for the Adam optimizer
by: Dereich, Steffen, et al.
Published: (2025)