PADAM: Parallel averaged Adam reduces the error for stochastic optimization in scientific machine learning
Fuente:
arXiv
Salvato in:
| Autori principali: | Jentzen, Arnulf, Kranz, Julian, Riekert, Adrian |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Averaged Adam accelerates stochastic optimization in the training of deep neural network approximations for partial differential equation and optimal control problems
di: Dereich, Steffen, et al.
Pubblicazione: (2025)
di: Dereich, Steffen, et al.
Pubblicazione: (2025)
Learning rate adaptive stochastic gradient descent optimization methods: numerical simulations for deep learning methods for partial differential equations and convergence analyses
di: Dereich, Steffen, et al.
Pubblicazione: (2024)
di: Dereich, Steffen, et al.
Pubblicazione: (2024)
Non-convergence to global minimizers for Adam and stochastic gradient descent optimization and constructions of local minimizers in the training of artificial neural networks
di: Jentzen, Arnulf, et al.
Pubblicazione: (2024)
di: Jentzen, Arnulf, et al.
Pubblicazione: (2024)
Convergence rates for the Adam optimizer
di: Dereich, Steffen, et al.
Pubblicazione: (2024)
di: Dereich, Steffen, et al.
Pubblicazione: (2024)
Algorithmically Designed Artificial Neural Networks (ADANNs): Higher order deep operator learning for parametric partial differential equations
di: Jentzen, Arnulf, et al.
Pubblicazione: (2023)
di: Jentzen, Arnulf, et al.
Pubblicazione: (2023)
Non-convergence to the optimal risk for Adam and stochastic gradient descent optimization in the training of deep neural networks
di: Do, Thang, et al.
Pubblicazione: (2025)
di: Do, Thang, et al.
Pubblicazione: (2025)
On the existence of minimizers in shallow residual ReLU neural network optimization landscapes
di: Dereich, Steffen, et al.
Pubblicazione: (2023)
di: Dereich, Steffen, et al.
Pubblicazione: (2023)
Real-time optimal control of high-dimensional parametrized systems by deep learning-based reduced order models
di: Tomasetto, Matteo, et al.
Pubblicazione: (2024)
di: Tomasetto, Matteo, et al.
Pubblicazione: (2024)
Non-convergence to global minimizers in data driven supervised deep learning: Adam and stochastic gradient descent optimization provably fail to converge to global minimizers in the training of deep neural networks with ReLU activation
di: Do, Thang, et al.
Pubblicazione: (2024)
di: Do, Thang, et al.
Pubblicazione: (2024)
Towards Quantifying the Preconditioning Effect of Adam
di: Das, Rudrajit, et al.
Pubblicazione: (2024)
di: Das, Rudrajit, et al.
Pubblicazione: (2024)
First-order methods for stochastic and finite-sum convex optimization with deterministic constraints
di: Lu, Zhaosong, et al.
Pubblicazione: (2025)
di: Lu, Zhaosong, et al.
Pubblicazione: (2025)
Sharp higher order convergence rates for the Adam optimizer
di: Dereich, Steffen, et al.
Pubblicazione: (2025)
di: Dereich, Steffen, et al.
Pubblicazione: (2025)
Flattened one-bit stochastic gradient descent: compressed distributed optimization with controlled variance
di: Stollenwerk, Alexander, et al.
Pubblicazione: (2024)
di: Stollenwerk, Alexander, et al.
Pubblicazione: (2024)
ODE approximation for the Adam algorithm: General and overparametrized setting
di: Dereich, Steffen, et al.
Pubblicazione: (2025)
di: Dereich, Steffen, et al.
Pubblicazione: (2025)
Modeling AdaGrad, RMSProp, and Adam with Integro-Differential Equations
di: Heredia, Carlos
Pubblicazione: (2024)
di: Heredia, Carlos
Pubblicazione: (2024)
Convergence of Adam for Non-convex Objectives: Relaxed Hyperparameters and Non-ergodic Case
di: He, Meixuan, et al.
Pubblicazione: (2023)
di: He, Meixuan, et al.
Pubblicazione: (2023)
UAdam: Unified Adam-Type Algorithmic Framework for Non-Convex Stochastic Optimization
di: Jiang, Yiming, et al.
Pubblicazione: (2023)
di: Jiang, Yiming, et al.
Pubblicazione: (2023)
Langevin dynamics based algorithm e-TH$\varepsilon$O POULA for stochastic optimization problems with discontinuous stochastic gradient
di: Lim, Dong-Young, et al.
Pubblicazione: (2022)
di: Lim, Dong-Young, et al.
Pubblicazione: (2022)
Latent feedback control of distributed systems in multiple scenarios through deep learning-based reduced order models
di: Tomasetto, Matteo, et al.
Pubblicazione: (2024)
di: Tomasetto, Matteo, et al.
Pubblicazione: (2024)
From Adam to Adam-Like Lagrangians: Second-Order Nonlocal Dynamics
di: Heredia, Carlos
Pubblicazione: (2026)
di: Heredia, Carlos
Pubblicazione: (2026)
An Overview on Machine Learning Methods for Partial Differential Equations: from Physics Informed Neural Networks to Deep Operator Learning
di: Gonon, Lukas, et al.
Pubblicazione: (2024)
di: Gonon, Lukas, et al.
Pubblicazione: (2024)
Cubic regularized subspace Newton for non-convex optimization
di: Zhao, Jim, et al.
Pubblicazione: (2024)
di: Zhao, Jim, et al.
Pubblicazione: (2024)
On bounds for norms of reparameterized ReLU artificial neural network parameters: sums of fractional powers of the Lipschitz norm control the network parameter vector
di: Jentzen, Arnulf, et al.
Pubblicazione: (2022)
di: Jentzen, Arnulf, et al.
Pubblicazione: (2022)
A single-loop SPIDER-type stochastic subgradient method for expectation-constrained nonconvex nonsmooth optimization
di: Liu, Wei, et al.
Pubblicazione: (2025)
di: Liu, Wei, et al.
Pubblicazione: (2025)
Continuum-marginal optimal transport: a mesh-free kernel method
di: Nakano, Yumiharu
Pubblicazione: (2026)
di: Nakano, Yumiharu
Pubblicazione: (2026)
A distributed semismooth Newton based augmented Lagrangian method for distributed optimization
di: Ma, Qihao, et al.
Pubblicazione: (2026)
di: Ma, Qihao, et al.
Pubblicazione: (2026)
A note on continuous-time online learning
di: Ying, Lexing
Pubblicazione: (2024)
di: Ying, Lexing
Pubblicazione: (2024)
Natural Riemannian gradient for learning functional tensor networks
di: Klug, Nikolas, et al.
Pubblicazione: (2026)
di: Klug, Nikolas, et al.
Pubblicazione: (2026)
Neural incomplete factorization: learning preconditioners for the conjugate gradient method
di: Häusner, Paul, et al.
Pubblicazione: (2023)
di: Häusner, Paul, et al.
Pubblicazione: (2023)
Non-asymptotic convergence analysis of the stochastic gradient Hamiltonian Monte Carlo algorithm with discontinuous stochastic gradient with applications to training of ReLU neural networks
di: Liang, Luxu, et al.
Pubblicazione: (2024)
di: Liang, Luxu, et al.
Pubblicazione: (2024)
Deep learning based numerical approximation algorithms for stochastic partial differential equations
di: Beck, Christian, et al.
Pubblicazione: (2020)
di: Beck, Christian, et al.
Pubblicazione: (2020)
Convergence rates for gradient descent in the training of overparameterized artificial neural networks with piecewise affine activation
di: Jentzen, Arnulf, et al.
Pubblicazione: (2021)
di: Jentzen, Arnulf, et al.
Pubblicazione: (2021)
Asymptotic stability properties and a priori bounds for Adam and other gradient descent optimization methods
di: Dereich, Steffen, et al.
Pubblicazione: (2025)
di: Dereich, Steffen, et al.
Pubblicazione: (2025)
Uniform a priori bounds and error analysis for the Adam stochastic gradient descent optimization method
di: Dereich, Steffen, et al.
Pubblicazione: (2026)
di: Dereich, Steffen, et al.
Pubblicazione: (2026)
On the convergence of stochastic variance reduced gradient for linear inverse problems
di: Jin, Bangti, et al.
Pubblicazione: (2025)
di: Jin, Bangti, et al.
Pubblicazione: (2025)
Martingale deep learning for very high dimensional quasi-linear partial differential equations and stochastic optimal controls
di: Cai, Wei, et al.
Pubblicazione: (2024)
di: Cai, Wei, et al.
Pubblicazione: (2024)
Parallel-in-iteration optimization using multigrid reduction-in-time
di: Araújo, G. H. M., et al.
Pubblicazione: (2026)
di: Araújo, G. H. M., et al.
Pubblicazione: (2026)
Convergence to good non-optimal critical points in the training of neural networks: Gradient descent optimization with one random initialization overcomes all bad non-global local minima with high probability
di: Ibragimov, Shokhrukh, et al.
Pubblicazione: (2022)
di: Ibragimov, Shokhrukh, et al.
Pubblicazione: (2022)
Space-time shape optimization of rotating electric machines
di: Cesarano, Alessio, et al.
Pubblicazione: (2024)
di: Cesarano, Alessio, et al.
Pubblicazione: (2024)
Creating walls to avoid unwanted points in root finding and optimization
di: Truong, Tuyen Trung
Pubblicazione: (2023)
di: Truong, Tuyen Trung
Pubblicazione: (2023)
Documenti analoghi
-
Averaged Adam accelerates stochastic optimization in the training of deep neural network approximations for partial differential equation and optimal control problems
di: Dereich, Steffen, et al.
Pubblicazione: (2025) -
Learning rate adaptive stochastic gradient descent optimization methods: numerical simulations for deep learning methods for partial differential equations and convergence analyses
di: Dereich, Steffen, et al.
Pubblicazione: (2024) -
Non-convergence to global minimizers for Adam and stochastic gradient descent optimization and constructions of local minimizers in the training of artificial neural networks
di: Jentzen, Arnulf, et al.
Pubblicazione: (2024) -
Convergence rates for the Adam optimizer
di: Dereich, Steffen, et al.
Pubblicazione: (2024) -
Algorithmically Designed Artificial Neural Networks (ADANNs): Higher order deep operator learning for parametric partial differential equations
di: Jentzen, Arnulf, et al.
Pubblicazione: (2023)