Clipping Improves Adam-Norm and AdaGrad-Norm when the Noise Is Heavy-Tailed
Fuente:
arXiv
Saved in:
| Main Authors: | Chezhegov, Savelii, Klyukin, Yaroslav, Semenov, Andrei, Beznosikov, Aleksandr, Gasnikov, Alexander, Horváth, Samuel, Takáč, Martin, Gorbunov, Eduard |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Convergence of Clipped-SGD for Convex $(L_0,L_1)$-Smooth Optimization with Heavy-Tailed Noise
by: Chezhegov, Savelii, et al.
Published: (2025)
by: Chezhegov, Savelii, et al.
Published: (2025)
Local SGD for Near-Quadratic Problems: Improving Convergence under Unconstrained Noise Conditions
by: Sadchikov, Andrey, et al.
Published: (2024)
by: Sadchikov, Andrey, et al.
Published: (2024)
Last Iterate Convergence of AdaGrad-Norm for Convex Non-Smooth Optimization
by: Preobrazhenskaia, Margarita, et al.
Published: (2026)
by: Preobrazhenskaia, Margarita, et al.
Published: (2026)
Remove that Square Root: A New Efficient Scale-Invariant Version of AdaGrad
by: Choudhury, Sayantan, et al.
Published: (2024)
by: Choudhury, Sayantan, et al.
Published: (2024)
Differentially Private Clipped-SGD: High-Probability Convergence with Arbitrary Clipping Level
by: Khah, Saleh Vatan, et al.
Published: (2025)
by: Khah, Saleh Vatan, et al.
Published: (2025)
A Riemannian AdaGrad-Norm Method
by: Bento, Glaydston de C., et al.
Published: (2025)
by: Bento, Glaydston de C., et al.
Published: (2025)
Accelerated Stochastic Gradient Method with Applications to Consensus Problem in Markov-Varying Networks
by: Solodkin, Vladimir, et al.
Published: (2024)
by: Solodkin, Vladimir, et al.
Published: (2024)
Incorporating Preconditioning into Accelerated Approaches: Theoretical Guarantees and Practical Improvement
by: Trifonov, Stepan, et al.
Published: (2025)
by: Trifonov, Stepan, et al.
Published: (2025)
Enhancing Stability of Physics-Informed Neural Network Training Through Saddle-Point Reformulation
by: Bylinkin, Dmitry, et al.
Published: (2025)
by: Bylinkin, Dmitry, et al.
Published: (2025)
Breaking the Heavy-Tailed Noise Barrier in Stochastic Optimization Problems
by: Puchkin, Nikita, et al.
Published: (2023)
by: Puchkin, Nikita, et al.
Published: (2023)
Preconditioned Norms: A Unified Framework for Steepest Descent, Quasi-Newton and Adaptive Methods
by: Veprikov, Andrey, et al.
Published: (2025)
by: Veprikov, Andrey, et al.
Published: (2025)
Modeling AdaGrad, RMSProp, and Adam with Integro-Differential Equations
by: Heredia, Carlos
Published: (2024)
by: Heredia, Carlos
Published: (2024)
Local Methods with Adaptivity via Scaling
by: Chezhegov, Savelii, et al.
Published: (2024)
by: Chezhegov, Savelii, et al.
Published: (2024)
Can Adaptive Gradient Methods Converge under Heavy-Tailed Noise? A Case Study of AdaGrad
by: Liu, Zijian
Published: (2026)
by: Liu, Zijian
Published: (2026)
Variance Reduction Methods Do Not Need to Compute Full Gradients: Improved Efficiency through Shuffling
by: Medyakov, Daniil, et al.
Published: (2025)
by: Medyakov, Daniil, et al.
Published: (2025)
Linear Convergence Rate in Convex Setup is Possible! Gradient Descent Method Variants under $(L_0,L_1)$-Smoothness
by: Lobanov, Aleksandr, et al.
Published: (2024)
by: Lobanov, Aleksandr, et al.
Published: (2024)
Similarity, Compression and Local Steps: Three Pillars of Efficient Communications for Distributed Variational Inequalities
by: Beznosikov, Aleksandr, et al.
Published: (2023)
by: Beznosikov, Aleksandr, et al.
Published: (2023)
Decentralized Finite-Sum Optimization over Time-Varying Networks
by: Metelev, Dmitry, et al.
Published: (2024)
by: Metelev, Dmitry, et al.
Published: (2024)
High-Probability Convergence for Composite and Distributed Stochastic Minimization and Variational Inequalities with Heavy-Tailed Noise
by: Gorbunov, Eduard, et al.
Published: (2023)
by: Gorbunov, Eduard, et al.
Published: (2023)
Sign Operator for Coping with Heavy-Tailed Noise in Non-Convex Optimization: High Probability Bounds Under $(L_0, L_1)$-Smoothness
by: Kornilov, Nikita, et al.
Published: (2025)
by: Kornilov, Nikita, et al.
Published: (2025)
AdaGrad under Anisotropic Smoothness
by: Liu, Yuxing, et al.
Published: (2024)
by: Liu, Yuxing, et al.
Published: (2024)
Adaptive Regularized Newton Method with Inexact Hessian
by: Shestakov, Aleksandr, et al.
Published: (2025)
by: Shestakov, Aleksandr, et al.
Published: (2025)
Sparse Concept Bottleneck Models: Gumbel Tricks in Contrastive Learning
by: Semenov, Andrei, et al.
Published: (2024)
by: Semenov, Andrei, et al.
Published: (2024)
High Probability Complexity Bounds for Non-Smooth Stochastic Optimization with Heavy-Tailed Noise
by: Gorbunov, Eduard, et al.
Published: (2021)
by: Gorbunov, Eduard, et al.
Published: (2021)
Revisiting Convergence of AdaGrad with Relaxed Assumptions
by: Hong, Yusu, et al.
Published: (2024)
by: Hong, Yusu, et al.
Published: (2024)
Median Clipping for Zeroth-order Non-Smooth Convex Optimization and Multi-Armed Bandit Problem with Heavy-tailed Symmetric Noise
by: Kornilov, Nikita, et al.
Published: (2024)
by: Kornilov, Nikita, et al.
Published: (2024)
A unified convergence theory for adaptive first-order methods in the nonconvex case, including AdaNorm, full and diagonal AdaGrad, Shampoo and Muo
by: Gratton, S., et al.
Published: (2026)
by: Gratton, S., et al.
Published: (2026)
FRUGAL: Memory-Efficient Optimization by Reducing State Overhead for Scalable Training
by: Zmushko, Philip, et al.
Published: (2024)
by: Zmushko, Philip, et al.
Published: (2024)
Where Does Warm-Up Come From? Adaptive Scheduling for Norm-Constrained Optimizers
by: Riabinin, Artem, et al.
Published: (2026)
by: Riabinin, Artem, et al.
Published: (2026)
Federated Learning Can Find Friends That Are Advantageous
by: Tupitsa, Nazarii, et al.
Published: (2024)
by: Tupitsa, Nazarii, et al.
Published: (2024)
Byzantine-Robust Optimization under $(L_0, L_1)$-Smoothness
by: Bolatov, Arman, et al.
Published: (2026)
by: Bolatov, Arman, et al.
Published: (2026)
Methods for Convex $(L_0,L_1)$-Smooth Optimization: Clipping, Acceleration, and Adaptivity
by: Gorbunov, Eduard, et al.
Published: (2024)
by: Gorbunov, Eduard, et al.
Published: (2024)
Bregman Proximal Method for Efficient Communications under Similarity
by: Beznosikov, Aleksandr, et al.
Published: (2023)
by: Beznosikov, Aleksandr, et al.
Published: (2023)
AdaGrad Meets Muon: Adaptive Stepsizes for Orthogonal Updates
by: Zhang, Minxin, et al.
Published: (2025)
by: Zhang, Minxin, et al.
Published: (2025)
Randomized gradient-free methods in convex optimization
by: Gasnikov, Alexander, et al.
Published: (2022)
by: Gasnikov, Alexander, et al.
Published: (2022)
Byzantine Robustness and Partial Participation Can Be Achieved at Once: Just Clip Gradient Differences
by: Malinovsky, Grigory, et al.
Published: (2023)
by: Malinovsky, Grigory, et al.
Published: (2023)
Activations and Gradients Compression for Model-Parallel Training
by: Rudakov, Mikhail, et al.
Published: (2024)
by: Rudakov, Mikhail, et al.
Published: (2024)
AdaBatchGrad: Combining Adaptive Batch Size and Adaptive Step Size
by: Ostroukhov, Petr, et al.
Published: (2024)
by: Ostroukhov, Petr, et al.
Published: (2024)
Random-reshuffled SARAH does not need a full gradient computations
by: Beznosikov, Aleksandr, et al.
Published: (2021)
by: Beznosikov, Aleksandr, et al.
Published: (2021)
AdaGrad-Diff: A New Version of the Adaptive Gradient Algorithm
by: Bojovic, Matia, et al.
Published: (2026)
by: Bojovic, Matia, et al.
Published: (2026)
Similar Items
-
Convergence of Clipped-SGD for Convex $(L_0,L_1)$-Smooth Optimization with Heavy-Tailed Noise
by: Chezhegov, Savelii, et al.
Published: (2025) -
Local SGD for Near-Quadratic Problems: Improving Convergence under Unconstrained Noise Conditions
by: Sadchikov, Andrey, et al.
Published: (2024) -
Last Iterate Convergence of AdaGrad-Norm for Convex Non-Smooth Optimization
by: Preobrazhenskaia, Margarita, et al.
Published: (2026) -
Remove that Square Root: A New Efficient Scale-Invariant Version of AdaGrad
by: Choudhury, Sayantan, et al.
Published: (2024) -
Differentially Private Clipped-SGD: High-Probability Convergence with Arbitrary Clipping Level
by: Khah, Saleh Vatan, et al.
Published: (2025)