Convergence of Steepest Descent and Adam under Non-Uniform Smoothness
Fuente:
arXiv
Saved in:
| Main Authors: | Vaswani, Sharan, Sun, Yifan, Babanezhad, Reza |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Armijo Line-search Can Make (Stochastic) Gradient Descent Provably Faster
by: Vaswani, Sharan, et al.
Published: (2025)
by: Vaswani, Sharan, et al.
Published: (2025)
Towards Noise-adaptive, Problem-adaptive (Accelerated) Stochastic Gradient Descent
by: Vaswani, Sharan, et al.
Published: (2021)
by: Vaswani, Sharan, et al.
Published: (2021)
(Accelerated) Noise-adaptive Stochastic Heavy-Ball Momentum
by: Dang, Anh, et al.
Published: (2024)
by: Dang, Anh, et al.
Published: (2024)
On the Convergence of Adam under Non-uniform Smoothness: Separability from SGDM and Beyond
by: Wang, Bohan, et al.
Published: (2024)
by: Wang, Bohan, et al.
Published: (2024)
Faster Acceleration for Steepest Descent
by: Bai, Cedar Site, et al.
Published: (2024)
by: Bai, Cedar Site, et al.
Published: (2024)
Glocal Smoothness: Line search and adaptive step sizes can help in theory too!
by: Fox, Curtis, et al.
Published: (2025)
by: Fox, Curtis, et al.
Published: (2025)
Provable Adaptivity of Adam under Non-uniform Smoothness
by: Wang, Bohan, et al.
Published: (2022)
by: Wang, Bohan, et al.
Published: (2022)
On the Convergence of Adam-Type Algorithm for Bilevel Optimization under Unbounded Smoothness
by: Gong, Xiaochuan, et al.
Published: (2025)
by: Gong, Xiaochuan, et al.
Published: (2025)
From Inverse Optimization to Feasibility to ERM
by: Mishra, Saurabh, et al.
Published: (2024)
by: Mishra, Saurabh, et al.
Published: (2024)
Adam-SHANG: A Convergent Adam-Type Method for Stochastic Smooth Convex Optimization
by: Yu, Yaxin, et al.
Published: (2026)
by: Yu, Yaxin, et al.
Published: (2026)
Preconditioned Norms: A Unified Framework for Steepest Descent, Quasi-Newton and Adaptive Methods
by: Veprikov, Andrey, et al.
Published: (2025)
by: Veprikov, Andrey, et al.
Published: (2025)
On Convergence of Adam for Stochastic Optimization under Relaxed Assumptions
by: Hong, Yusu, et al.
Published: (2024)
by: Hong, Yusu, et al.
Published: (2024)
Implicit Bias and Convergence of Matrix Stochastic Mirror Descent
by: Akhtiamov, Danil, et al.
Published: (2026)
by: Akhtiamov, Danil, et al.
Published: (2026)
Convergence of Spectral Descent for Non-smooth Optimization
by: Yang, Yixuan, et al.
Published: (2026)
by: Yang, Yixuan, et al.
Published: (2026)
The Rich and the Simple: On the Implicit Bias of Adam and SGD
by: Vasudeva, Bhavya, et al.
Published: (2025)
by: Vasudeva, Bhavya, et al.
Published: (2025)
On the Inherent Privacy of Zeroth Order Projected Gradient Descent
by: Gupta, Devansh, et al.
Published: (2025)
by: Gupta, Devansh, et al.
Published: (2025)
MGDA Converges under Generalized Smoothness, Provably
by: Zhang, Qi, et al.
Published: (2024)
by: Zhang, Qi, et al.
Published: (2024)
Dissecting Discrete Soft Actor-Critic: Limitations and Principled Alternatives
by: Asad, Reza, et al.
Published: (2025)
by: Asad, Reza, et al.
Published: (2025)
A Mirror Descent Perspective of Smoothed Sign Descent
by: Wang, Shuyang, et al.
Published: (2024)
by: Wang, Shuyang, et al.
Published: (2024)
Adam-HNAG: A Convergent Reformulation of Adam with Accelerated Rate
by: Yu, Yaxin, et al.
Published: (2026)
by: Yu, Yaxin, et al.
Published: (2026)
Adam Converges Without Any Modification On Update Rules
by: Zhang, Yushun, et al.
Published: (2026)
by: Zhang, Yushun, et al.
Published: (2026)
Faster Convergence of Stochastic Accelerated Gradient Descent under Interpolation
by: Mishkin, Aaron, et al.
Published: (2024)
by: Mishkin, Aaron, et al.
Published: (2024)
On Convergence of Incremental Gradient for Non-Convex Smooth Functions
by: Koloskova, Anastasia, et al.
Published: (2023)
by: Koloskova, Anastasia, et al.
Published: (2023)
Convergence rates for the Adam optimizer
by: Dereich, Steffen, et al.
Published: (2024)
by: Dereich, Steffen, et al.
Published: (2024)
AltGDmin: Alternating GD and Minimization for Partly-Decoupled (Federated) Optimization
by: Vaswani, Namrata
Published: (2025)
by: Vaswani, Namrata
Published: (2025)
Convergence Guarantees for RMSProp and Adam in Generalized-smooth Non-convex Optimization with Affine Noise Variance
by: Zhang, Qi, et al.
Published: (2024)
by: Zhang, Qi, et al.
Published: (2024)
The Method of Infinite Descent
by: Batley, Reza T., et al.
Published: (2025)
by: Batley, Reza T., et al.
Published: (2025)
Tight Lower Bounds under Asymmetric High-Order Hölder Smoothness and Uniform Convexity
by: Bai, Cedar Site, et al.
Published: (2024)
by: Bai, Cedar Site, et al.
Published: (2024)
Adam-family Methods for Nonsmooth Optimization with Convergence Guarantees
by: Xiao, Nachuan, et al.
Published: (2023)
by: Xiao, Nachuan, et al.
Published: (2023)
A Theoretical and Empirical Study on the Convergence of Adam with an "Exact" Constant Step Size in Non-Convex Settings
by: Mazumder, Alokendu, et al.
Published: (2023)
by: Mazumder, Alokendu, et al.
Published: (2023)
Convergence of Adam for Non-convex Objectives: Relaxed Hyperparameters and Non-ergodic Case
by: He, Meixuan, et al.
Published: (2023)
by: He, Meixuan, et al.
Published: (2023)
Provably Convergent Decentralized Optimization over Directed Graphs under Generalized Smoothness
by: Bo, Yanan, et al.
Published: (2026)
by: Bo, Yanan, et al.
Published: (2026)
Bilevel Optimization under Unbounded Smoothness: A New Algorithm and Convergence Analysis
by: Hao, Jie, et al.
Published: (2024)
by: Hao, Jie, et al.
Published: (2024)
Learning Provably Improves the Convergence of Gradient Descent
by: Song, Qingyu, et al.
Published: (2025)
by: Song, Qingyu, et al.
Published: (2025)
On the Convergence of Policy in Unregularized Policy Mirror Descent
by: Lin, Dachao, et al.
Published: (2022)
by: Lin, Dachao, et al.
Published: (2022)
Convergence of Alternating Gradient Descent for Matrix Factorization
by: Ward, Rachel, et al.
Published: (2023)
by: Ward, Rachel, et al.
Published: (2023)
SPGD: Steepest Perturbed Gradient Descent Optimization
by: Vahedi, Amir M., et al.
Published: (2024)
by: Vahedi, Amir M., et al.
Published: (2024)
Quantitative Convergence Analysis of Projected Stochastic Gradient Descent for Non-Convex Losses via the Goldstein Subdifferential
by: Zheng, Yuping, et al.
Published: (2025)
by: Zheng, Yuping, et al.
Published: (2025)
Convergence Analysis of Stochastic Gradient Descent with MCMC Estimators
by: Li, Tianyou, et al.
Published: (2023)
by: Li, Tianyou, et al.
Published: (2023)
On the Convergence of Policy Mirror Descent with Temporal Difference Evaluation
by: Liu, Jiacai, et al.
Published: (2025)
by: Liu, Jiacai, et al.
Published: (2025)
Similar Items
-
Armijo Line-search Can Make (Stochastic) Gradient Descent Provably Faster
by: Vaswani, Sharan, et al.
Published: (2025) -
Towards Noise-adaptive, Problem-adaptive (Accelerated) Stochastic Gradient Descent
by: Vaswani, Sharan, et al.
Published: (2021) -
(Accelerated) Noise-adaptive Stochastic Heavy-Ball Momentum
by: Dang, Anh, et al.
Published: (2024) -
On the Convergence of Adam under Non-uniform Smoothness: Separability from SGDM and Beyond
by: Wang, Bohan, et al.
Published: (2024) -
Faster Acceleration for Steepest Descent
by: Bai, Cedar Site, et al.
Published: (2024)