Adam with model exponential moving average is effective for nonconvex optimization
Fuente:
arXiv
Saved in:
| Main Authors: | Ahn, Kwangjun, Cutkosky, Ashok |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
General framework for online-to-nonconvex conversion: Schedule-free SGD is also effective for nonconvex optimization
by: Ahn, Kwangjun, et al.
Published: (2024)
by: Ahn, Kwangjun, et al.
Published: (2024)
Understanding Adam Optimizer via Online Learning of Updates: Adam is FTRL in Disguise
by: Ahn, Kwangjun, et al.
Published: (2024)
by: Ahn, Kwangjun, et al.
Published: (2024)
Random Scaling and Momentum for Non-smooth Non-convex Optimization
by: Zhang, Qinzi, et al.
Published: (2024)
by: Zhang, Qinzi, et al.
Published: (2024)
Fully Unconstrained Online Learning
by: Cutkosky, Ashok, et al.
Published: (2024)
by: Cutkosky, Ashok, et al.
Published: (2024)
Parameter-free Mirror Descent
by: Jacobsen, Andrew, et al.
Published: (2022)
by: Jacobsen, Andrew, et al.
Published: (2022)
Unconstrained Robust Online Convex Optimization
by: Zhang, Jiujia, et al.
Published: (2025)
by: Zhang, Jiujia, et al.
Published: (2025)
Does SGD really happen in tiny subspaces?
by: Song, Minhak, et al.
Published: (2024)
by: Song, Minhak, et al.
Published: (2024)
Reevaluating Theoretical Analysis Methods for Optimization in Deep Learning
by: Tran, Hoang, et al.
Published: (2024)
by: Tran, Hoang, et al.
Published: (2024)
Optimal Stochastic Non-smooth Non-convex Optimization through Online-to-Non-convex Conversion
by: Cutkosky, Ashok, et al.
Published: (2023)
by: Cutkosky, Ashok, et al.
Published: (2023)
How to escape sharp minima with random perturbations
by: Ahn, Kwangjun, et al.
Published: (2023)
by: Ahn, Kwangjun, et al.
Published: (2023)
Private Zeroth-Order Nonsmooth Nonconvex Optimization
by: Zhang, Qinzi, et al.
Published: (2024)
by: Zhang, Qinzi, et al.
Published: (2024)
Linear attention is (maybe) all you need (to understand transformer optimization)
by: Ahn, Kwangjun, et al.
Published: (2023)
by: Ahn, Kwangjun, et al.
Published: (2023)
Through the River: Understanding the Benefit of Schedule-Free Methods for Language Model Training
by: Song, Minhak, et al.
Published: (2025)
by: Song, Minhak, et al.
Published: (2025)
The inexact power augmented Lagrangian method for constrained nonconvex optimization
by: Bodard, Alexander, et al.
Published: (2024)
by: Bodard, Alexander, et al.
Published: (2024)
Newton-CG methods for nonconvex unconstrained optimization with Hölder continuous Hessian
by: He, Chuan, et al.
Published: (2023)
by: He, Chuan, et al.
Published: (2023)
Block majorization-minimization with diminishing radius for constrained nonsmooth nonconvex optimization
by: Lyu, Hanbaek, et al.
Published: (2020)
by: Lyu, Hanbaek, et al.
Published: (2020)
PADAM: Parallel averaged Adam reduces the error for stochastic optimization in scientific machine learning
by: Jentzen, Arnulf, et al.
Published: (2025)
by: Jentzen, Arnulf, et al.
Published: (2025)
Convergence rates for the Adam optimizer
by: Dereich, Steffen, et al.
Published: (2024)
by: Dereich, Steffen, et al.
Published: (2024)
Projected gradient methods for nonconvex and stochastic smooth optimization: new complexities and auto-conditioned stepsizes
by: Lan, Guanghui, et al.
Published: (2024)
by: Lan, Guanghui, et al.
Published: (2024)
Task-optimal data-driven surrogate models for eNMPC via differentiable simulation and optimization
by: Mayfrank, Daniel, et al.
Published: (2024)
by: Mayfrank, Daniel, et al.
Published: (2024)
A stochastic smoothing framework for nonconvex-nonconcave min-sum-max problems with applications to Wasserstein distributionally robust optimization
by: Liu, Wei, et al.
Published: (2025)
by: Liu, Wei, et al.
Published: (2025)
Nonasymptotic analysis of Stochastic Gradient Hamiltonian Monte Carlo under local conditions for nonconvex optimization
by: Akyildiz, Ömer Deniz, et al.
Published: (2020)
by: Akyildiz, Ömer Deniz, et al.
Published: (2020)
Avoiding strict saddle points of nonconvex regularized problems
by: Bai, Luwei, et al.
Published: (2024)
by: Bai, Luwei, et al.
Published: (2024)
The Road Less Scheduled
by: Defazio, Aaron, et al.
Published: (2024)
by: Defazio, Aaron, et al.
Published: (2024)
HomeAdam: Adam and AdamW Algorithms Sometimes Go Home to Obtain Better Provable Generalization
by: Huang, Feihu, et al.
Published: (2026)
by: Huang, Feihu, et al.
Published: (2026)
Adam-HNAG: A Convergent Reformulation of Adam with Accelerated Rate
by: Yu, Yaxin, et al.
Published: (2026)
by: Yu, Yaxin, et al.
Published: (2026)
Adam on Local Time: Addressing Nonstationarity in RL with Relative Adam Timesteps
by: Ellis, Benjamin, et al.
Published: (2024)
by: Ellis, Benjamin, et al.
Published: (2024)
On exploration of an interior mirror descent flow for stochastic nonconvex constrained problem
by: Ding, Kuangyu, et al.
Published: (2025)
by: Ding, Kuangyu, et al.
Published: (2025)
A randomized algorithm for nonconvex minimization with inexact evaluations and complexity guarantees
by: Li, Shuyao, et al.
Published: (2023)
by: Li, Shuyao, et al.
Published: (2023)
Instance-optimal stochastic convex optimization: Can we improve upon sample-average and robust stochastic approximation?
by: Jiang, Liwei, et al.
Published: (2026)
by: Jiang, Liwei, et al.
Published: (2026)
Dion: Distributed Orthonormalized Updates
by: Ahn, Kwangjun, et al.
Published: (2025)
by: Ahn, Kwangjun, et al.
Published: (2025)
Convergence of SGD with momentum in the nonconvex case: A time window-based analysis
by: Qiu, Junwen, et al.
Published: (2024)
by: Qiu, Junwen, et al.
Published: (2024)
Adam-SHANG: A Convergent Adam-Type Method for Stochastic Smooth Convex Optimization
by: Yu, Yaxin, et al.
Published: (2026)
by: Yu, Yaxin, et al.
Published: (2026)
Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling
by: Srećković, Teodora, et al.
Published: (2025)
by: Srećković, Teodora, et al.
Published: (2025)
Non-convergence to global minimizers for Adam and stochastic gradient descent optimization and constructions of local minimizers in the training of artificial neural networks
by: Jentzen, Arnulf, et al.
Published: (2024)
by: Jentzen, Arnulf, et al.
Published: (2024)
From exponential to finite/fixed-time stability: Applications to optimization
by: Ozaslan, Ibrahim K., et al.
Published: (2024)
by: Ozaslan, Ibrahim K., et al.
Published: (2024)
An accelerated first-order regularized momentum descent ascent algorithm for stochastic nonconvex-concave minimax problems
by: Zhang, Huiling, et al.
Published: (2023)
by: Zhang, Huiling, et al.
Published: (2023)
The Rich and the Simple: On the Implicit Bias of Adam and SGD
by: Vasudeva, Bhavya, et al.
Published: (2025)
by: Vasudeva, Bhavya, et al.
Published: (2025)
Finite sample learning of moving targets
by: Vertovec, Nikolaus, et al.
Published: (2024)
by: Vertovec, Nikolaus, et al.
Published: (2024)
On Convergence of Adam for Stochastic Optimization under Relaxed Assumptions
by: Hong, Yusu, et al.
Published: (2024)
by: Hong, Yusu, et al.
Published: (2024)
Similar Items
-
General framework for online-to-nonconvex conversion: Schedule-free SGD is also effective for nonconvex optimization
by: Ahn, Kwangjun, et al.
Published: (2024) -
Understanding Adam Optimizer via Online Learning of Updates: Adam is FTRL in Disguise
by: Ahn, Kwangjun, et al.
Published: (2024) -
Random Scaling and Momentum for Non-smooth Non-convex Optimization
by: Zhang, Qinzi, et al.
Published: (2024) -
Fully Unconstrained Online Learning
by: Cutkosky, Ashok, et al.
Published: (2024) -
Parameter-free Mirror Descent
by: Jacobsen, Andrew, et al.
Published: (2022)