Heavy-Tail Phenomenon in Decentralized SGD
Fuente:
arXiv
Saved in:
| Main Authors: | Gurbuzbalaban, Mert, Hu, Yuanhan, Simsekli, Umut, Yuan, Kun, Zhu, Lingjiong |
|---|---|
| Format: | Preprint |
| Published: |
2022
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Algorithmic Stability of Stochastic Gradient Descent with Momentum under Heavy-Tailed Noise
by: Dang, Thanh, et al.
Published: (2025)
by: Dang, Thanh, et al.
Published: (2025)
Privacy of SGD under Gaussian or Heavy-Tailed Noise: Guarantees without Gradient Clipping
by: Şimşekli, Umut, et al.
Published: (2024)
by: Şimşekli, Umut, et al.
Published: (2024)
DIGing--SGLD: Decentralized and Scalable Langevin Sampling over Time--Varying Networks
by: Bajwa, Waheed U., et al.
Published: (2025)
by: Bajwa, Waheed U., et al.
Published: (2025)
Generalized EXTRA stochastic gradient Langevin dynamics
by: Gurbuzbalaban, Mert, et al.
Published: (2024)
by: Gurbuzbalaban, Mert, et al.
Published: (2024)
Rényi Differential Privacy for Heavy-Tailed SDEs via Fractional Poincaré Inequalities
by: Dupuis, Benjamin, et al.
Published: (2025)
by: Dupuis, Benjamin, et al.
Published: (2025)
Penalized Overdamped and Underdamped Langevin Monte Carlo Algorithms for Constrained Sampling
by: Gürbüzbalaban, Mert, et al.
Published: (2022)
by: Gürbüzbalaban, Mert, et al.
Published: (2022)
Revisiting Gradient Normalization and Clipping for Nonconvex SGD under Heavy-Tailed Noise: Necessity, Sufficiency, and Acceleration
by: Sun, Tao, et al.
Published: (2024)
by: Sun, Tao, et al.
Published: (2024)
RESIST: Resilient Decentralized Learning Using Consensus Gradient Descent
by: Fang, Cheng, et al.
Published: (2025)
by: Fang, Cheng, et al.
Published: (2025)
Can SGD Handle Heavy-Tailed Noise?
by: Fatkhullin, Ilyas, et al.
Published: (2025)
by: Fatkhullin, Ilyas, et al.
Published: (2025)
From Gradient Clipping to Normalization for Heavy Tailed SGD
by: Hübler, Florian, et al.
Published: (2024)
by: Hübler, Florian, et al.
Published: (2024)
Accelerated Gradient Methods with Biased Gradient Estimates: Risk Sensitivity, High-Probability Guarantees, and Large Deviation Bounds
by: Gürbüzbalaban, Mert, et al.
Published: (2025)
by: Gürbüzbalaban, Mert, et al.
Published: (2025)
Sharp High-Probability Rates for Nonlinear SGD under Heavy-Tailed Noise via Symmetrization
by: Armacki, Aleksandar, et al.
Published: (2025)
by: Armacki, Aleksandar, et al.
Published: (2025)
Convergence of Clipped-SGD for Convex $(L_0,L_1)$-Smooth Optimization with Heavy-Tailed Noise
by: Chezhegov, Savelii, et al.
Published: (2025)
by: Chezhegov, Savelii, et al.
Published: (2025)
Near-Optimal Decentralized Stochastic Nonconvex Optimization with Heavy-Tailed Noise
by: Wang, Menglian, et al.
Published: (2026)
by: Wang, Menglian, et al.
Published: (2026)
The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training
by: Schaipp, Fabian, et al.
Published: (2025)
by: Schaipp, Fabian, et al.
Published: (2025)
From PowerSGD to PowerSGD+: Low-Rank Gradient Compression for Distributed Optimization with Convergence Guarantees
by: Xie, Shengping, et al.
Published: (2025)
by: Xie, Shengping, et al.
Published: (2025)
Implicit Compressibility of Overparametrized Neural Networks Trained with Heavy-Tailed SGD
by: Wan, Yijun, et al.
Published: (2023)
by: Wan, Yijun, et al.
Published: (2023)
Stochastic Weakly Convex Optimization Under Heavy-Tailed Noises
by: Zhu, Tianxi, et al.
Published: (2025)
by: Zhu, Tianxi, et al.
Published: (2025)
An Accelerated Primal Dual Algorithm with Backtracking for Decentralized Constrained Optimization
by: Xu, Qiushui, et al.
Published: (2025)
by: Xu, Qiushui, et al.
Published: (2025)
Tight Long-Term Tail Decay of (Clipped) SGD in Non-Convex Optimization
by: Armacki, Aleksandar, et al.
Published: (2026)
by: Armacki, Aleksandar, et al.
Published: (2026)
Why is Normalization Preferred? A Worst-Case Complexity Theory for Stochastically Preconditioned SGD under Heavy-Tailed Noise
by: Fang, Yuchen, et al.
Published: (2026)
by: Fang, Yuchen, et al.
Published: (2026)
High-Probability Convergence Guarantees of Decentralized SGD
by: Armacki, Aleksandar, et al.
Published: (2025)
by: Armacki, Aleksandar, et al.
Published: (2025)
Asynchronous Decentralized SGD under Non-Convexity: A Block-Coordinate Descent Framework
by: Zhou, Yijie, et al.
Published: (2025)
by: Zhou, Yijie, et al.
Published: (2025)
Decentralized Bilevel Optimization: A Perspective from Transient Iteration Complexity
by: Kong, Boao, et al.
Published: (2024)
by: Kong, Boao, et al.
Published: (2024)
SPARKLE: A Unified Single-Loop Primal-Dual Framework for Decentralized Bilevel Optimization
by: Zhu, Shuchen, et al.
Published: (2024)
by: Zhu, Shuchen, et al.
Published: (2024)
Efficient Private SCO for Heavy-Tailed Data via Averaged Clipping
by: Jin, Chenhan, et al.
Published: (2022)
by: Jin, Chenhan, et al.
Published: (2022)
Optimal Asynchronous Stochastic Nonconvex Optimization under Heavy-Tailed Noise
by: Wu, Yidong, et al.
Published: (2026)
by: Wu, Yidong, et al.
Published: (2026)
In-Expectation Convergence of Stochastic Gradient Methods under Heavy-Tailed Noise
by: Liu, Zijian
Published: (2026)
by: Liu, Zijian
Published: (2026)
Generalization Bounds for Heavy-Tailed SDEs through the Fractional Fokker-Planck Equation
by: Dupuis, Benjamin, et al.
Published: (2024)
by: Dupuis, Benjamin, et al.
Published: (2024)
Tracking the Median of Gradients with a Stochastic Proximal Point Method
by: Schaipp, Fabian, et al.
Published: (2024)
by: Schaipp, Fabian, et al.
Published: (2024)
Online Convex Optimization with Heavy Tails: Old Algorithms, New Regrets, and Applications
by: Liu, Zijian
Published: (2025)
by: Liu, Zijian
Published: (2025)
Scale-Invariant Neural Network Optimization: Norm Geometry and Heavy-Tailed Noise
by: Zhang, Jiayu, et al.
Published: (2026)
by: Zhang, Jiayu, et al.
Published: (2026)
Muon with Nesterov Momentum: Heavy-Tailed Noise and (Randomized) Inexact Polar Decomposition
by: Choudhury, Sayantan, et al.
Published: (2026)
by: Choudhury, Sayantan, et al.
Published: (2026)
High Probability Complexity Bounds for Non-Smooth Stochastic Optimization with Heavy-Tailed Noise
by: Gorbunov, Eduard, et al.
Published: (2021)
by: Gorbunov, Eduard, et al.
Published: (2021)
Clipping Improves Adam-Norm and AdaGrad-Norm when the Noise Is Heavy-Tailed
by: Chezhegov, Savelii, et al.
Published: (2024)
by: Chezhegov, Savelii, et al.
Published: (2024)
SLowcal-SGD: Slow Query Points Improve Local-SGD for Stochastic Convex Optimization
by: Dahan, Tehila, et al.
Published: (2023)
by: Dahan, Tehila, et al.
Published: (2023)
A Minibatch-SGD-Based Learning Meta-Policy for Inventory Systems with Myopic Optimal Policy
by: Lyu, Jiameng, et al.
Published: (2024)
by: Lyu, Jiameng, et al.
Published: (2024)
Making SGD Parameter-Free
by: Carmon, Yair, et al.
Published: (2022)
by: Carmon, Yair, et al.
Published: (2022)
On the Trajectories of SGD Without Replacement
by: Beneventano, Pierfrancesco
Published: (2023)
by: Beneventano, Pierfrancesco
Published: (2023)
Muon Converges under Heavy-Tailed Noise: Nonconvex Hölder-Smooth Empirical Risk Minimization
by: Iiduka, Hideaki
Published: (2026)
by: Iiduka, Hideaki
Published: (2026)
Similar Items
-
Algorithmic Stability of Stochastic Gradient Descent with Momentum under Heavy-Tailed Noise
by: Dang, Thanh, et al.
Published: (2025) -
Privacy of SGD under Gaussian or Heavy-Tailed Noise: Guarantees without Gradient Clipping
by: Şimşekli, Umut, et al.
Published: (2024) -
DIGing--SGLD: Decentralized and Scalable Langevin Sampling over Time--Varying Networks
by: Bajwa, Waheed U., et al.
Published: (2025) -
Generalized EXTRA stochastic gradient Langevin dynamics
by: Gurbuzbalaban, Mert, et al.
Published: (2024) -
Rényi Differential Privacy for Heavy-Tailed SDEs via Fractional Poincaré Inequalities
by: Dupuis, Benjamin, et al.
Published: (2025)