Revisiting Convergence of AdaGrad with Relaxed Assumptions
Fuente:
arXiv
Saved in:
| Main Authors: | Hong, Yusu, Lin, Junhong |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
On Convergence of Adam for Stochastic Optimization under Relaxed Assumptions
by: Hong, Yusu, et al.
Published: (2024)
by: Hong, Yusu, et al.
Published: (2024)
AdaGrad under Anisotropic Smoothness
by: Liu, Yuxing, et al.
Published: (2024)
by: Liu, Yuxing, et al.
Published: (2024)
AdaGrad Meets Muon: Adaptive Stepsizes for Orthogonal Updates
by: Zhang, Minxin, et al.
Published: (2025)
by: Zhang, Minxin, et al.
Published: (2025)
AdaGrad-Diff: A New Version of the Adaptive Gradient Algorithm
by: Bojovic, Matia, et al.
Published: (2026)
by: Bojovic, Matia, et al.
Published: (2026)
Can Adaptive Gradient Methods Converge under Heavy-Tailed Noise? A Case Study of AdaGrad
by: Liu, Zijian
Published: (2026)
by: Liu, Zijian
Published: (2026)
Modeling AdaGrad, RMSProp, and Adam with Integro-Differential Equations
by: Heredia, Carlos
Published: (2024)
by: Heredia, Carlos
Published: (2024)
Clipping Improves Adam-Norm and AdaGrad-Norm when the Noise Is Heavy-Tailed
by: Chezhegov, Savelii, et al.
Published: (2024)
by: Chezhegov, Savelii, et al.
Published: (2024)
Remove that Square Root: A New Efficient Scale-Invariant Version of AdaGrad
by: Choudhury, Sayantan, et al.
Published: (2024)
by: Choudhury, Sayantan, et al.
Published: (2024)
Provable Complexity Improvement of AdaGrad over SGD: Upper and Lower Bounds in Stochastic Non-Convex Optimization
by: Jiang, Ruichen, et al.
Published: (2024)
by: Jiang, Ruichen, et al.
Published: (2024)
A Riemannian AdaGrad-Norm Method
by: Bento, Glaydston de C., et al.
Published: (2025)
by: Bento, Glaydston de C., et al.
Published: (2025)
Last Iterate Convergence of AdaGrad-Norm for Convex Non-Smooth Optimization
by: Preobrazhenskaia, Margarita, et al.
Published: (2026)
by: Preobrazhenskaia, Margarita, et al.
Published: (2026)
Convergence Analysis of Stochastic Accelerated Gradient Methods for Generalized Smooth Optimizations
by: Yu, Chenhao, et al.
Published: (2025)
by: Yu, Chenhao, et al.
Published: (2025)
Universality of AdaGrad Stepsizes for Stochastic Optimization: Inexact Oracle, Acceleration and Variance Reduction
by: Rodomanov, Anton, et al.
Published: (2024)
by: Rodomanov, Anton, et al.
Published: (2024)
AdAdaGrad: Adaptive Batch Size Schemes for Adaptive Gradient Methods
by: Lau, Tim Tsz-Kit, et al.
Published: (2024)
by: Lau, Tim Tsz-Kit, et al.
Published: (2024)
AdaBatchGrad: Combining Adaptive Batch Size and Adaptive Step Size
by: Ostroukhov, Petr, et al.
Published: (2024)
by: Ostroukhov, Petr, et al.
Published: (2024)
Stability and convergence analysis of AdaGrad for non-convex optimization via novel stopping time-based techniques
by: Jin, Ruinan, et al.
Published: (2024)
by: Jin, Ruinan, et al.
Published: (2024)
Learning Over-Relaxation Policies for ADMM with Convergence Guarantees
by: Lin, Junan, et al.
Published: (2026)
by: Lin, Junan, et al.
Published: (2026)
Revisiting Convergence: Shuffling Complexity Beyond Lipschitz Smoothness
by: He, Qi, et al.
Published: (2025)
by: He, Qi, et al.
Published: (2025)
Revisiting the Last-Iterate Convergence of Stochastic Gradient Methods
by: Liu, Zijian, et al.
Published: (2023)
by: Liu, Zijian, et al.
Published: (2023)
Learning Provably Improves the Convergence of Gradient Descent
by: Song, Qingyu, et al.
Published: (2025)
by: Song, Qingyu, et al.
Published: (2025)
Revisiting Subgradient Method: Complexity and Convergence Beyond Lipschitz Continuity
by: Li, Xiao, et al.
Published: (2023)
by: Li, Xiao, et al.
Published: (2023)
GradPower: Powering Gradients for Faster Language Model Pre-Training
by: Wang, Jinbo, et al.
Published: (2025)
by: Wang, Jinbo, et al.
Published: (2025)
Towards Weaker Variance Assumptions for Stochastic Optimization
by: Alacaoglu, Ahmet, et al.
Published: (2025)
by: Alacaoglu, Ahmet, et al.
Published: (2025)
Dual Acceleration for Minimax Optimization: Linear Convergence Under Relaxed Assumptions
by: Li, Jingwang, et al.
Published: (2025)
by: Li, Jingwang, et al.
Published: (2025)
Adaptive Variance Reduction for Stochastic Optimization under Weaker Assumptions
by: Jiang, Wei, et al.
Published: (2024)
by: Jiang, Wei, et al.
Published: (2024)
Why Smooth Stability Assumptions Fail for ReLU Learning
by: Katende, Ronald
Published: (2025)
by: Katende, Ronald
Published: (2025)
Solving Stochastic Variational Inequalities without the Bounded Variance Assumption
by: Alacaoglu, Ahmet, et al.
Published: (2026)
by: Alacaoglu, Ahmet, et al.
Published: (2026)
PolarGrad: A Class of Matrix-Gradient Optimizers from a Unifying Preconditioning Perspective
by: Lau, Tim Tsz-Kit, et al.
Published: (2025)
by: Lau, Tim Tsz-Kit, et al.
Published: (2025)
Convergence Rate Analysis of LION
by: Dong, Yiming, et al.
Published: (2024)
by: Dong, Yiming, et al.
Published: (2024)
Reusing Historical Trajectories in Natural Policy Gradient via Importance Sampling: Convergence and Convergence Rate
by: Lin, Yifan, et al.
Published: (2024)
by: Lin, Yifan, et al.
Published: (2024)
Convergence of Adam for Non-convex Objectives: Relaxed Hyperparameters and Non-ergodic Case
by: He, Meixuan, et al.
Published: (2023)
by: He, Meixuan, et al.
Published: (2023)
AdaFisher: Adaptive Second Order Optimization via Fisher Information
by: Gomes, Damien Martins, et al.
Published: (2024)
by: Gomes, Damien Martins, et al.
Published: (2024)
First Provable Guarantees for Practical Private FL: Beyond Restrictive Assumptions
by: Shulgin, Egor, et al.
Published: (2025)
by: Shulgin, Egor, et al.
Published: (2025)
A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies
by: Nanda, Phalguni, et al.
Published: (2025)
by: Nanda, Phalguni, et al.
Published: (2025)
On the Convergence of Policy in Unregularized Policy Mirror Descent
by: Lin, Dachao, et al.
Published: (2022)
by: Lin, Dachao, et al.
Published: (2022)
GeoAdaLer: Geometric Insights into Adaptive Stochastic Gradient Descent Algorithms
by: Eleh, Chinedu, et al.
Published: (2024)
by: Eleh, Chinedu, et al.
Published: (2024)
TiAda: A Time-scale Adaptive Algorithm for Nonconvex Minimax Optimization
by: Li, Xiang, et al.
Published: (2022)
by: Li, Xiang, et al.
Published: (2022)
AdaSwitch: An Adaptive Switching Meta-Algorithm for Learning-Augmented Bounded-Influence Problems
by: Chen, Xi, et al.
Published: (2025)
by: Chen, Xi, et al.
Published: (2025)
Semidefinite Relaxations of the Gromov-Wasserstein Distance
by: Chen, Junyu, et al.
Published: (2023)
by: Chen, Junyu, et al.
Published: (2023)
Achieving $ε^{-2}$ Sample Complexity for Single-Loop Actor-Critic under Minimal Assumptions
by: Hamza, Ishaq, et al.
Published: (2026)
by: Hamza, Ishaq, et al.
Published: (2026)
Similar Items
-
On Convergence of Adam for Stochastic Optimization under Relaxed Assumptions
by: Hong, Yusu, et al.
Published: (2024) -
AdaGrad under Anisotropic Smoothness
by: Liu, Yuxing, et al.
Published: (2024) -
AdaGrad Meets Muon: Adaptive Stepsizes for Orthogonal Updates
by: Zhang, Minxin, et al.
Published: (2025) -
AdaGrad-Diff: A New Version of the Adaptive Gradient Algorithm
by: Bojovic, Matia, et al.
Published: (2026) -
Can Adaptive Gradient Methods Converge under Heavy-Tailed Noise? A Case Study of AdaGrad
by: Liu, Zijian
Published: (2026)