Modeling AdaGrad, RMSProp, and Adam with Integro-Differential Equations
Fuente:
arXiv
Saved in:
| Main Author: | Heredia, Carlos |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
From Adam to Adam-Like Lagrangians: Second-Order Nonlocal Dynamics
by: Heredia, Carlos
Published: (2026)
by: Heredia, Carlos
Published: (2026)
AdaGrad under Anisotropic Smoothness
by: Liu, Yuxing, et al.
Published: (2024)
by: Liu, Yuxing, et al.
Published: (2024)
Clipping Improves Adam-Norm and AdaGrad-Norm when the Noise Is Heavy-Tailed
by: Chezhegov, Savelii, et al.
Published: (2024)
by: Chezhegov, Savelii, et al.
Published: (2024)
Revisiting Convergence of AdaGrad with Relaxed Assumptions
by: Hong, Yusu, et al.
Published: (2024)
by: Hong, Yusu, et al.
Published: (2024)
AdaGrad Meets Muon: Adaptive Stepsizes for Orthogonal Updates
by: Zhang, Minxin, et al.
Published: (2025)
by: Zhang, Minxin, et al.
Published: (2025)
AdaGrad-Diff: A New Version of the Adaptive Gradient Algorithm
by: Bojovic, Matia, et al.
Published: (2026)
by: Bojovic, Matia, et al.
Published: (2026)
Towards Quantifying the Preconditioning Effect of Adam
by: Das, Rudrajit, et al.
Published: (2024)
by: Das, Rudrajit, et al.
Published: (2024)
Control, Optimal Transport and Neural Differential Equations in Supervised Learning
by: Phung, Minh-Nhat, et al.
Published: (2025)
by: Phung, Minh-Nhat, et al.
Published: (2025)
Convergence of Adam for Non-convex Objectives: Relaxed Hyperparameters and Non-ergodic Case
by: He, Meixuan, et al.
Published: (2023)
by: He, Meixuan, et al.
Published: (2023)
UAdam: Unified Adam-Type Algorithmic Framework for Non-Convex Stochastic Optimization
by: Jiang, Yiming, et al.
Published: (2023)
by: Jiang, Yiming, et al.
Published: (2023)
PADAM: Parallel averaged Adam reduces the error for stochastic optimization in scientific machine learning
by: Jentzen, Arnulf, et al.
Published: (2025)
by: Jentzen, Arnulf, et al.
Published: (2025)
A Natural Primal-Dual Hybrid Gradient Method for Adversarial Neural Network Training on Solving Partial Differential Equations
by: Liu, Shu, et al.
Published: (2024)
by: Liu, Shu, et al.
Published: (2024)
Remove that Square Root: A New Efficient Scale-Invariant Version of AdaGrad
by: Choudhury, Sayantan, et al.
Published: (2024)
by: Choudhury, Sayantan, et al.
Published: (2024)
Averaged Adam accelerates stochastic optimization in the training of deep neural network approximations for partial differential equation and optimal control problems
by: Dereich, Steffen, et al.
Published: (2025)
by: Dereich, Steffen, et al.
Published: (2025)
Provable Complexity Improvement of AdaGrad over SGD: Upper and Lower Bounds in Stochastic Non-Convex Optimization
by: Jiang, Ruichen, et al.
Published: (2024)
by: Jiang, Ruichen, et al.
Published: (2024)
Can Adaptive Gradient Methods Converge under Heavy-Tailed Noise? A Case Study of AdaGrad
by: Liu, Zijian
Published: (2026)
by: Liu, Zijian
Published: (2026)
Stochastic Langevin Differential Inclusions with Applications to Machine Learning
by: Difonzo, Fabio V., et al.
Published: (2022)
by: Difonzo, Fabio V., et al.
Published: (2022)
A Riemannian AdaGrad-Norm Method
by: Bento, Glaydston de C., et al.
Published: (2025)
by: Bento, Glaydston de C., et al.
Published: (2025)
Subhomogeneous Deep Equilibrium Models
by: Sittoni, Pietro, et al.
Published: (2024)
by: Sittoni, Pietro, et al.
Published: (2024)
Multi-level Optimal Control with Neural Surrogate Models
by: Kalise, Dante, et al.
Published: (2024)
by: Kalise, Dante, et al.
Published: (2024)
Riemannian Preconditioned LoRA for Fine-Tuning Foundation Models
by: Zhang, Fangzhao, et al.
Published: (2024)
by: Zhang, Fangzhao, et al.
Published: (2024)
Efficient Algorithms for Regularized Nonnegative Scale-invariant Low-rank Approximation Models
by: Cohen, Jeremy E., et al.
Published: (2024)
by: Cohen, Jeremy E., et al.
Published: (2024)
WinQ: Accelerating Quantization-Aware Training of Language Models Around Saddle Points
by: Li, Dongyue, et al.
Published: (2026)
by: Li, Dongyue, et al.
Published: (2026)
Convergence Guarantees for RMSProp and Adam in Generalized-smooth Non-convex Optimization with Affine Noise Variance
by: Zhang, Qi, et al.
Published: (2024)
by: Zhang, Qi, et al.
Published: (2024)
Learning rate adaptive stochastic gradient descent optimization methods: numerical simulations for deep learning methods for partial differential equations and convergence analyses
by: Dereich, Steffen, et al.
Published: (2024)
by: Dereich, Steffen, et al.
Published: (2024)
On the numerical reliability of nonsmooth autodiff: a MaxPool case study
by: Boustany, Ryan
Published: (2024)
by: Boustany, Ryan
Published: (2024)
A Gauss-Newton Approach for Min-Max Optimization in Generative Adversarial Networks
by: Mishra, Neel, et al.
Published: (2024)
by: Mishra, Neel, et al.
Published: (2024)
Flattened one-bit stochastic gradient descent: compressed distributed optimization with controlled variance
by: Stollenwerk, Alexander, et al.
Published: (2024)
by: Stollenwerk, Alexander, et al.
Published: (2024)
Efficient Trajectory Inference in Wasserstein Space Using Consecutive Averaging
by: Banerjee, Amartya, et al.
Published: (2024)
by: Banerjee, Amartya, et al.
Published: (2024)
Anderson Acceleration in Nonsmooth Problems: Local Convergence via Active Manifold Identification
by: Li, Kexin, et al.
Published: (2024)
by: Li, Kexin, et al.
Published: (2024)
KANtrol: A Physics-Informed Kolmogorov-Arnold Network Framework for Solving Multi-Dimensional and Fractional Optimal Control Problems
by: Aghaei, Alireza Afzal
Published: (2024)
by: Aghaei, Alireza Afzal
Published: (2024)
Real-time optimal control of high-dimensional parametrized systems by deep learning-based reduced order models
by: Tomasetto, Matteo, et al.
Published: (2024)
by: Tomasetto, Matteo, et al.
Published: (2024)
Cubic regularized subspace Newton for non-convex optimization
by: Zhao, Jim, et al.
Published: (2024)
by: Zhao, Jim, et al.
Published: (2024)
Super Gradient Descent: Global Optimization requires Global Gradient
by: Achour, Seifeddine
Published: (2024)
by: Achour, Seifeddine
Published: (2024)
Latent feedback control of distributed systems in multiple scenarios through deep learning-based reduced order models
by: Tomasetto, Matteo, et al.
Published: (2024)
by: Tomasetto, Matteo, et al.
Published: (2024)
Quantitative Convergences of Lie Group Momentum Optimizers
by: Kong, Lingkai, et al.
Published: (2024)
by: Kong, Lingkai, et al.
Published: (2024)
Learning incomplete factorization preconditioners for GMRES
by: Häusner, Paul, et al.
Published: (2024)
by: Häusner, Paul, et al.
Published: (2024)
A note on continuous-time online learning
by: Ying, Lexing
Published: (2024)
by: Ying, Lexing
Published: (2024)
Symmetry & Critical Points
by: Arjevani, Yossi
Published: (2024)
by: Arjevani, Yossi
Published: (2024)
ADMM for Structured Fractional Minimization
by: Yuan, Ganzhao
Published: (2024)
by: Yuan, Ganzhao
Published: (2024)
Similar Items
-
From Adam to Adam-Like Lagrangians: Second-Order Nonlocal Dynamics
by: Heredia, Carlos
Published: (2026) -
AdaGrad under Anisotropic Smoothness
by: Liu, Yuxing, et al.
Published: (2024) -
Clipping Improves Adam-Norm and AdaGrad-Norm when the Noise Is Heavy-Tailed
by: Chezhegov, Savelii, et al.
Published: (2024) -
Revisiting Convergence of AdaGrad with Relaxed Assumptions
by: Hong, Yusu, et al.
Published: (2024) -
AdaGrad Meets Muon: Adaptive Stepsizes for Orthogonal Updates
by: Zhang, Minxin, et al.
Published: (2025)