Parameter-free Clipped Gradient Descent Meets Polyak
Fuente:
arXiv
Saved in:
| Main Authors: | Takezawa, Yuki, Bao, Han, Sato, Ryoma, Niwa, Kenta, Yamada, Makoto |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Delayed Momentum Aggregation: Communication-efficient Byzantine-robust Federated Learning with Partial Participation
by: Otsuka, Kaoru, et al.
Published: (2025)
by: Otsuka, Kaoru, et al.
Published: (2025)
Necessary and Sufficient Watermark for Large Language Models
by: Takezawa, Yuki, et al.
Published: (2023)
by: Takezawa, Yuki, et al.
Published: (2023)
A Local Polyak-Lojasiewicz and Descent Lemma of Gradient Descent For Overparametrized Linear Models
by: Xu, Ziqing, et al.
Published: (2025)
by: Xu, Ziqing, et al.
Published: (2025)
Gradient Descent with Polyak's Momentum Finds Flatter Minima via Large Catapults
by: Phunyaphibarn, Prin, et al.
Published: (2023)
by: Phunyaphibarn, Prin, et al.
Published: (2023)
Scalable Decentralized Learning with Teleportation
by: Takezawa, Yuki, et al.
Published: (2025)
by: Takezawa, Yuki, et al.
Published: (2025)
Parameter-free Mirror Descent
by: Jacobsen, Andrew, et al.
Published: (2022)
by: Jacobsen, Andrew, et al.
Published: (2022)
Momentum Does Not Reduce Stochastic Noise in Stochastic Gradient Descent
by: Sato, Naoki, et al.
Published: (2024)
by: Sato, Naoki, et al.
Published: (2024)
On the Convergence of the Gradient Descent Method with Stochastic Fixed-point Rounding Errors under the Polyak-Lojasiewicz Inequality
by: Xia, Lu, et al.
Published: (2023)
by: Xia, Lu, et al.
Published: (2023)
Parameter Symmetry and Noise Equilibrium of Stochastic Gradient Descent
by: Ziyin, Liu, et al.
Published: (2024)
by: Ziyin, Liu, et al.
Published: (2024)
Gradient Clipping Beyond Vector Norms: A Spectral Approach for Matrix-Valued Parameters
by: Yukhimchuk, Alexander, et al.
Published: (2026)
by: Yukhimchuk, Alexander, et al.
Published: (2026)
Using Stochastic Gradient Descent to Smooth Nonconvex Functions: Analysis of Implicit Graduated Optimization
by: Sato, Naoki, et al.
Published: (2023)
by: Sato, Naoki, et al.
Published: (2023)
Unraveling the Gradient Descent Dynamics of Transformers
by: Song, Bingqing, et al.
Published: (2024)
by: Song, Bingqing, et al.
Published: (2024)
On Penalty-based Bilevel Gradient Descent Method
by: Shen, Han, et al.
Published: (2023)
by: Shen, Han, et al.
Published: (2023)
FedMuon: Federated Learning with Bias-corrected LMO-based Optimization
by: Takezawa, Yuki, et al.
Published: (2025)
by: Takezawa, Yuki, et al.
Published: (2025)
Exploiting Similarity for Computation and Communication-Efficient Decentralized Optimization
by: Takezawa, Yuki, et al.
Published: (2025)
by: Takezawa, Yuki, et al.
Published: (2025)
DoWG Unleashed: An Efficient Universal Parameter-Free Gradient Descent Method
by: Khaled, Ahmed, et al.
Published: (2023)
by: Khaled, Ahmed, et al.
Published: (2023)
From Gradient Clipping to Normalization for Heavy Tailed SGD
by: Hübler, Florian, et al.
Published: (2024)
by: Hübler, Florian, et al.
Published: (2024)
Stochastic Adaptive Gradient Descent Without Descent
by: Aujol, Jean-François, et al.
Published: (2025)
by: Aujol, Jean-François, et al.
Published: (2025)
Relationship between Batch Size and Number of Steps Needed for Nonconvex Optimization of Stochastic Gradient Descent using Armijo Line Search
by: Tsukada, Yuki, et al.
Published: (2023)
by: Tsukada, Yuki, et al.
Published: (2023)
Corner Gradient Descent
by: Yarotsky, Dmitry
Published: (2025)
by: Yarotsky, Dmitry
Published: (2025)
Adaptive Conditional Gradient Descent
by: Khademi, Abbas, et al.
Published: (2025)
by: Khademi, Abbas, et al.
Published: (2025)
$k$-SVD with Gradient Descent
by: Jedra, Yassir, et al.
Published: (2025)
by: Jedra, Yassir, et al.
Published: (2025)
Anytime Acceleration of Gradient Descent
by: Zhang, Zihan, et al.
Published: (2024)
by: Zhang, Zihan, et al.
Published: (2024)
Regularized Gradient Clipping Provably Trains Wide and Deep Neural Networks
by: Tucat, Matteo, et al.
Published: (2024)
by: Tucat, Matteo, et al.
Published: (2024)
Safeguarded Stochastic Polyak Step Sizes for Non-smooth Optimization: Robust Performance Without Small (Sub)Gradients
by: Oikonomou, Dimitris, et al.
Published: (2025)
by: Oikonomou, Dimitris, et al.
Published: (2025)
Stochastic Gradient Descent with Adaptive Data
by: Che, Ethan, et al.
Published: (2024)
by: Che, Ethan, et al.
Published: (2024)
Stochastic Gradient Descent with Strategic Querying
by: Jiang, Nanfei, et al.
Published: (2025)
by: Jiang, Nanfei, et al.
Published: (2025)
Low-Tubal-Rank Tensor Recovery via Factorized Gradient Descent
by: Liu, Zhiyu, et al.
Published: (2024)
by: Liu, Zhiyu, et al.
Published: (2024)
Mirror and Preconditioned Gradient Descent in Wasserstein Space
by: Bonet, Clément, et al.
Published: (2024)
by: Bonet, Clément, et al.
Published: (2024)
Derivatives of Stochastic Gradient Descent in parametric optimization
by: Iutzeler, Franck, et al.
Published: (2024)
by: Iutzeler, Franck, et al.
Published: (2024)
Learning Provably Improves the Convergence of Gradient Descent
by: Song, Qingyu, et al.
Published: (2025)
by: Song, Qingyu, et al.
Published: (2025)
Enhancing Fractional Gradient Descent with Learned Optimizers
by: Sobotka, Jan, et al.
Published: (2025)
by: Sobotka, Jan, et al.
Published: (2025)
Convergence of Alternating Gradient Descent for Matrix Factorization
by: Ward, Rachel, et al.
Published: (2023)
by: Ward, Rachel, et al.
Published: (2023)
Non-convex Stochastic Composite Optimization with Polyak Momentum
by: Gao, Yuan, et al.
Published: (2024)
by: Gao, Yuan, et al.
Published: (2024)
Constrained Online Convex Optimization with Polyak Feasibility Steps
by: Hutchinson, Spencer, et al.
Published: (2025)
by: Hutchinson, Spencer, et al.
Published: (2025)
Efficient Low-Tubal-Rank Tensor Estimation via Alternating Preconditioned Gradient Descent
by: Liu, Zhiyu, et al.
Published: (2025)
by: Liu, Zhiyu, et al.
Published: (2025)
Scaling Laws for Gradient Descent and Sign Descent for Linear Bigram Models under Zipf's Law
by: Kunstner, Frederik, et al.
Published: (2025)
by: Kunstner, Frederik, et al.
Published: (2025)
Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping
by: Liu, Zijian, et al.
Published: (2024)
by: Liu, Zijian, et al.
Published: (2024)
The Sample Complexity of Gradient Descent in Stochastic Convex Optimization
by: Livni, Roi
Published: (2024)
by: Livni, Roi
Published: (2024)
Open Problem: Anytime Convergence Rate of Gradient Descent
by: Kornowski, Guy, et al.
Published: (2024)
by: Kornowski, Guy, et al.
Published: (2024)
Similar Items
-
Delayed Momentum Aggregation: Communication-efficient Byzantine-robust Federated Learning with Partial Participation
by: Otsuka, Kaoru, et al.
Published: (2025) -
Necessary and Sufficient Watermark for Large Language Models
by: Takezawa, Yuki, et al.
Published: (2023) -
A Local Polyak-Lojasiewicz and Descent Lemma of Gradient Descent For Overparametrized Linear Models
by: Xu, Ziqing, et al.
Published: (2025) -
Gradient Descent with Polyak's Momentum Finds Flatter Minima via Large Catapults
by: Phunyaphibarn, Prin, et al.
Published: (2023) -
Scalable Decentralized Learning with Teleportation
by: Takezawa, Yuki, et al.
Published: (2025)