Convergence of Adam for Non-convex Objectives: Relaxed Hyperparameters and Non-ergodic Case
Fuente:
arXiv
Saved in:
| Main Authors: | He, Meixuan, Liang, Yuqing, Liu, Jinlan, Xu, Dongpo |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
UAdam: Unified Adam-Type Algorithmic Framework for Non-Convex Stochastic Optimization
by: Jiang, Yiming, et al.
Published: (2023)
by: Jiang, Yiming, et al.
Published: (2023)
Towards Quantifying the Preconditioning Effect of Adam
by: Das, Rudrajit, et al.
Published: (2024)
by: Das, Rudrajit, et al.
Published: (2024)
Modeling AdaGrad, RMSProp, and Adam with Integro-Differential Equations
by: Heredia, Carlos
Published: (2024)
by: Heredia, Carlos
Published: (2024)
On the Width Scaling of Neural Optimizers Under Matrix Operator Norms I: Row/Column Normalization and Hyperparameter Transfer
by: Xu, Ruihan, et al.
Published: (2026)
by: Xu, Ruihan, et al.
Published: (2026)
PADAM: Parallel averaged Adam reduces the error for stochastic optimization in scientific machine learning
by: Jentzen, Arnulf, et al.
Published: (2025)
by: Jentzen, Arnulf, et al.
Published: (2025)
Cubic regularized subspace Newton for non-convex optimization
by: Zhao, Jim, et al.
Published: (2024)
by: Zhao, Jim, et al.
Published: (2024)
From Adam to Adam-Like Lagrangians: Second-Order Nonlocal Dynamics
by: Heredia, Carlos
Published: (2026)
by: Heredia, Carlos
Published: (2026)
Averaged Adam accelerates stochastic optimization in the training of deep neural network approximations for partial differential equation and optimal control problems
by: Dereich, Steffen, et al.
Published: (2025)
by: Dereich, Steffen, et al.
Published: (2025)
First-order methods for stochastic and finite-sum convex optimization with deterministic constraints
by: Lu, Zhaosong, et al.
Published: (2025)
by: Lu, Zhaosong, et al.
Published: (2025)
Quantitative Convergences of Lie Group Momentum Optimizers
by: Kong, Lingkai, et al.
Published: (2024)
by: Kong, Lingkai, et al.
Published: (2024)
Linear convergence of forward-backward accelerated algorithms without knowledge of the modulus of strong convexity
by: Li, Bowen, et al.
Published: (2023)
by: Li, Bowen, et al.
Published: (2023)
Gradient is All You Need? How Consensus-Based Optimization can be Interpreted as a Stochastic Relaxation of Gradient Descent
by: Riedl, Konstantin, et al.
Published: (2023)
by: Riedl, Konstantin, et al.
Published: (2023)
Non-asymptotic convergence analysis of the stochastic gradient Hamiltonian Monte Carlo algorithm with discontinuous stochastic gradient with applications to training of ReLU neural networks
by: Liang, Luxu, et al.
Published: (2024)
by: Liang, Luxu, et al.
Published: (2024)
Solving the Offline and Online Min-Max Problem of Non-smooth Submodular-Concave Functions: A Zeroth-Order Approach
by: Farzin, Amir Ali, et al.
Published: (2026)
by: Farzin, Amir Ali, et al.
Published: (2026)
Last-Iterate Convergence of Randomized Kaczmarz and SGD with Greedy Step Size
by: Dereziński, Michał, et al.
Published: (2026)
by: Dereziński, Michał, et al.
Published: (2026)
Anderson Acceleration in Nonsmooth Problems: Local Convergence via Active Manifold Identification
by: Li, Kexin, et al.
Published: (2024)
by: Li, Kexin, et al.
Published: (2024)
Convergence of two-timescale gradient descent ascent dynamics: finite-dimensional and mean-field perspectives
by: An, Jing, et al.
Published: (2025)
by: An, Jing, et al.
Published: (2025)
On the Convergence of the Gradient Descent Method with Stochastic Fixed-point Rounding Errors under the Polyak-Lojasiewicz Inequality
by: Xia, Lu, et al.
Published: (2023)
by: Xia, Lu, et al.
Published: (2023)
Why is Normalization Preferred? A Worst-Case Complexity Theory for Stochastically Preconditioned SGD under Heavy-Tailed Noise
by: Fang, Yuchen, et al.
Published: (2026)
by: Fang, Yuchen, et al.
Published: (2026)
BC-ADMM: An Efficient Non-convex Constrained Optimizer with Robotic Applications
by: Pan, Zherong, et al.
Published: (2025)
by: Pan, Zherong, et al.
Published: (2025)
Affine-Invariant Global Non-Asymptotic Convergence Analysis of BFGS under Self-Concordance
by: Jin, Qiujiang, et al.
Published: (2025)
by: Jin, Qiujiang, et al.
Published: (2025)
A distributed semismooth Newton based augmented Lagrangian method for distributed optimization
by: Ma, Qihao, et al.
Published: (2026)
by: Ma, Qihao, et al.
Published: (2026)
A Natural Primal-Dual Hybrid Gradient Method for Adversarial Neural Network Training on Solving Partial Differential Equations
by: Liu, Shu, et al.
Published: (2024)
by: Liu, Shu, et al.
Published: (2024)
sparseGeoHOPCA: A Geometric Solution to Sparse Higher-Order PCA Without Covariance Estimation
by: Xu, Renjie, et al.
Published: (2025)
by: Xu, Renjie, et al.
Published: (2025)
Primal-Dual Methods for Nonsmooth Nonconvex Optimization with Orthogonality Constraints
by: Zhu, Linglingzhi, et al.
Published: (2026)
by: Zhu, Linglingzhi, et al.
Published: (2026)
Convergence Guarantees for RMSProp and Adam in Generalized-smooth Non-convex Optimization with Affine Noise Variance
by: Zhang, Qi, et al.
Published: (2024)
by: Zhang, Qi, et al.
Published: (2024)
WinQ: Accelerating Quantization-Aware Training of Language Models Around Saddle Points
by: Li, Dongyue, et al.
Published: (2026)
by: Li, Dongyue, et al.
Published: (2026)
Convergence Analysis of Fractional Gradient Descent
by: Aggarwal, Ashwani
Published: (2023)
by: Aggarwal, Ashwani
Published: (2023)
Convergence of Spectral Descent for Non-smooth Optimization
by: Yang, Yixuan, et al.
Published: (2026)
by: Yang, Yixuan, et al.
Published: (2026)
Randomized Kaczmarz Methods with Beyond-Krylov Convergence
by: Dereziński, Michał, et al.
Published: (2025)
by: Dereziński, Michał, et al.
Published: (2025)
Scalable Acceleration for Classification-Based Derivative-Free Optimization
by: Han, Tianyi, et al.
Published: (2023)
by: Han, Tianyi, et al.
Published: (2023)
Error Feedback Can Accurately Compress Preconditioners
by: Modoranu, Ionut-Vlad, et al.
Published: (2023)
by: Modoranu, Ionut-Vlad, et al.
Published: (2023)
Solving Elliptic Optimal Control Problems via Neural Networks and Optimality System
by: Dai, Yongcheng, et al.
Published: (2023)
by: Dai, Yongcheng, et al.
Published: (2023)
Enhanced Adaptive Gradient Algorithms for Nonconvex-PL Minimax Optimization
by: Huang, Feihu, et al.
Published: (2023)
by: Huang, Feihu, et al.
Published: (2023)
Neural incomplete factorization: learning preconditioners for the conjugate gradient method
by: Häusner, Paul, et al.
Published: (2023)
by: Häusner, Paul, et al.
Published: (2023)
Provably Faster Gradient Descent via Long Steps
by: Grimmer, Benjamin
Published: (2023)
by: Grimmer, Benjamin
Published: (2023)
Adaptive Proximal Gradient Method for Convex Optimization
by: Malitsky, Yura, et al.
Published: (2023)
by: Malitsky, Yura, et al.
Published: (2023)
A Block Coordinate Descent Method for Nonsmooth Composite Optimization under Orthogonality Constraints
by: Yuan, Ganzhao
Published: (2023)
by: Yuan, Ganzhao
Published: (2023)
Dynamic Proximal Gradient Algorithms for Schatten-$p$ Quasi-Norm Regularized Problems
by: Shen, Weiping, et al.
Published: (2026)
by: Shen, Weiping, et al.
Published: (2026)
Fundamental Bias in Inverting Random Sampling Matrices with Application to Sub-sampled Newton
by: Niu, Chengmei, et al.
Published: (2025)
by: Niu, Chengmei, et al.
Published: (2025)
Similar Items
-
UAdam: Unified Adam-Type Algorithmic Framework for Non-Convex Stochastic Optimization
by: Jiang, Yiming, et al.
Published: (2023) -
Towards Quantifying the Preconditioning Effect of Adam
by: Das, Rudrajit, et al.
Published: (2024) -
Modeling AdaGrad, RMSProp, and Adam with Integro-Differential Equations
by: Heredia, Carlos
Published: (2024) -
On the Width Scaling of Neural Optimizers Under Matrix Operator Norms I: Row/Column Normalization and Hyperparameter Transfer
by: Xu, Ruihan, et al.
Published: (2026) -
PADAM: Parallel averaged Adam reduces the error for stochastic optimization in scientific machine learning
by: Jentzen, Arnulf, et al.
Published: (2025)