Does Worst-Performing Agent Lead the Pack? Analyzing Agent Dynamics in Unified Distributed SGD
Fuente:
arXiv
Salvato in:
| Autori principali: | Hu, Jie, Ma, Yi-Ting, Eun, Do Young |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Accelerating Distributed Stochastic Optimization via Self-Repellent Random Walks
di: Hu, Jie, et al.
Pubblicazione: (2024)
di: Hu, Jie, et al.
Pubblicazione: (2024)
Central Limit Theorem for Two-Timescale Stochastic Approximation with Markovian Noise: Theory and Applications
di: Hu, Jie, et al.
Pubblicazione: (2024)
di: Hu, Jie, et al.
Pubblicazione: (2024)
Byzantine-Robust Distributed SGD: A Unified Analysis and Tight Error Bounds
di: Ruan, Boyuan, et al.
Pubblicazione: (2026)
di: Ruan, Boyuan, et al.
Pubblicazione: (2026)
Does SGD really happen in tiny subspaces?
di: Song, Minhak, et al.
Pubblicazione: (2024)
di: Song, Minhak, et al.
Pubblicazione: (2024)
Demystifying SGD with Doubly Stochastic Gradients
di: Kim, Kyurae, et al.
Pubblicazione: (2024)
di: Kim, Kyurae, et al.
Pubblicazione: (2024)
A Comprehensive Framework for Analyzing the Convergence of Adam: Bridging the Gap with SGD
di: Jin, Ruinan, et al.
Pubblicazione: (2024)
di: Jin, Ruinan, et al.
Pubblicazione: (2024)
Does SGD Seek Flatness or Sharpness? An Exactly Solvable Model
di: Xu, Yizhou, et al.
Pubblicazione: (2026)
di: Xu, Yizhou, et al.
Pubblicazione: (2026)
SGD with Adaptive Preconditioning: Unified Analysis and Momentum Acceleration
di: Kovalev, Dmitry
Pubblicazione: (2025)
di: Kovalev, Dmitry
Pubblicazione: (2025)
Scaling Laws of SignSGD in Linear Regression: When Does It Outperform SGD?
di: Kim, Jihwan, et al.
Pubblicazione: (2026)
di: Kim, Jihwan, et al.
Pubblicazione: (2026)
Shadowheart SGD: Distributed Asynchronous SGD with Optimal Time Complexity Under Arbitrary Computation and Communication Heterogeneity
di: Tyurin, Alexander, et al.
Pubblicazione: (2024)
di: Tyurin, Alexander, et al.
Pubblicazione: (2024)
From PowerSGD to PowerSGD+: Low-Rank Gradient Compression for Distributed Optimization with Convergence Guarantees
di: Xie, Shengping, et al.
Pubblicazione: (2025)
di: Xie, Shengping, et al.
Pubblicazione: (2025)
Non-Euclidean SGD for Structured Optimization: Unified Analysis and Improved Rates
di: Kovalev, Dmitry, et al.
Pubblicazione: (2025)
di: Kovalev, Dmitry, et al.
Pubblicazione: (2025)
Heavy-Tail Phenomenon in Decentralized SGD
di: Gurbuzbalaban, Mert, et al.
Pubblicazione: (2022)
di: Gurbuzbalaban, Mert, et al.
Pubblicazione: (2022)
Why is Normalization Preferred? A Worst-Case Complexity Theory for Stochastically Preconditioned SGD under Heavy-Tailed Noise
di: Fang, Yuchen, et al.
Pubblicazione: (2026)
di: Fang, Yuchen, et al.
Pubblicazione: (2026)
VAMO: Efficient Zeroth-Order Variance Reduction for SGD with Faster Convergence
di: Chen, Jiahe, et al.
Pubblicazione: (2025)
di: Chen, Jiahe, et al.
Pubblicazione: (2025)
Improving the Worst-Case Bidirectional Communication Complexity for Nonconvex Distributed Optimization under Function Similarity
di: Gruntkowska, Kaja, et al.
Pubblicazione: (2024)
di: Gruntkowska, Kaja, et al.
Pubblicazione: (2024)
SLowcal-SGD: Slow Query Points Improve Local-SGD for Stochastic Convex Optimization
di: Dahan, Tehila, et al.
Pubblicazione: (2023)
di: Dahan, Tehila, et al.
Pubblicazione: (2023)
SGD for Variational Inference: Tackling Unbounded Variance via Preconditioning and Dynamic Batching
di: Labarrière, Hippolyte, et al.
Pubblicazione: (2026)
di: Labarrière, Hippolyte, et al.
Pubblicazione: (2026)
Making SGD Parameter-Free
di: Carmon, Yair, et al.
Pubblicazione: (2022)
di: Carmon, Yair, et al.
Pubblicazione: (2022)
On the Trajectories of SGD Without Replacement
di: Beneventano, Pierfrancesco
Pubblicazione: (2023)
di: Beneventano, Pierfrancesco
Pubblicazione: (2023)
Convergence of SGD with momentum in the nonconvex case: A time window-based analysis
di: Qiu, Junwen, et al.
Pubblicazione: (2024)
di: Qiu, Junwen, et al.
Pubblicazione: (2024)
Dimension-adapted Momentum Outscales SGD
di: Ferbach, Damien, et al.
Pubblicazione: (2025)
di: Ferbach, Damien, et al.
Pubblicazione: (2025)
SGD with memory: fundamental properties and stochastic acceleration
di: Yarotsky, Dmitry, et al.
Pubblicazione: (2024)
di: Yarotsky, Dmitry, et al.
Pubblicazione: (2024)
The Rich and the Simple: On the Implicit Bias of Adam and SGD
di: Vasudeva, Bhavya, et al.
Pubblicazione: (2025)
di: Vasudeva, Bhavya, et al.
Pubblicazione: (2025)
Sign-SGD via Parameter-Free Optimization
di: Medyakov, Daniil, et al.
Pubblicazione: (2025)
di: Medyakov, Daniil, et al.
Pubblicazione: (2025)
Can SGD Handle Heavy-Tailed Noise?
di: Fatkhullin, Ilyas, et al.
Pubblicazione: (2025)
di: Fatkhullin, Ilyas, et al.
Pubblicazione: (2025)
A Unified Framework for Analyzing Meta-algorithms in Online Convex Optimization
di: Pedramfar, Mohammad, et al.
Pubblicazione: (2024)
di: Pedramfar, Mohammad, et al.
Pubblicazione: (2024)
Worst-case generation via minimax optimization in Wasserstein space
di: Cheng, Xiuyuan, et al.
Pubblicazione: (2025)
di: Cheng, Xiuyuan, et al.
Pubblicazione: (2025)
Logarithmically Quantized Distributed Optimization over Dynamic Multi-Agent Networks
di: Doostmohammadian, Mohammadreza, et al.
Pubblicazione: (2024)
di: Doostmohammadian, Mohammadreza, et al.
Pubblicazione: (2024)
Edge of Stochastic Stability: Revisiting the Edge of Stability for SGD
di: Andreyev, Arseniy, et al.
Pubblicazione: (2024)
di: Andreyev, Arseniy, et al.
Pubblicazione: (2024)
From Gradient Clipping to Normalization for Heavy Tailed SGD
di: Hübler, Florian, et al.
Pubblicazione: (2024)
di: Hübler, Florian, et al.
Pubblicazione: (2024)
The Optimality of (Accelerated) SGD for High-Dimensional Quadratic Optimization
di: Zhang, Haihan, et al.
Pubblicazione: (2024)
di: Zhang, Haihan, et al.
Pubblicazione: (2024)
Dual-Delayed Asynchronous SGD for Arbitrarily Heterogeneous Data
di: Wang, Xiaolu, et al.
Pubblicazione: (2024)
di: Wang, Xiaolu, et al.
Pubblicazione: (2024)
SGD with Partial Hessian for Deep Neural Networks Optimization
di: Sun, Ying, et al.
Pubblicazione: (2024)
di: Sun, Ying, et al.
Pubblicazione: (2024)
Faster Convergence of Local SGD for Over-Parameterized Models
di: Qin, Tiancheng, et al.
Pubblicazione: (2022)
di: Qin, Tiancheng, et al.
Pubblicazione: (2022)
The Marginal Value of Momentum for Small Learning Rate SGD
di: Wang, Runzhe, et al.
Pubblicazione: (2023)
di: Wang, Runzhe, et al.
Pubblicazione: (2023)
Phases of Muon: When Muon Eclipses SignSGD
di: Paquette, Elliot, et al.
Pubblicazione: (2026)
di: Paquette, Elliot, et al.
Pubblicazione: (2026)
Optimal Projection-Free Adaptive SGD for Matrix Optimization
di: Kovalev, Dmitry
Pubblicazione: (2026)
di: Kovalev, Dmitry
Pubblicazione: (2026)
Global Convergence of SGD On Two Layer Neural Nets
di: Gopalani, Pulkit, et al.
Pubblicazione: (2022)
di: Gopalani, Pulkit, et al.
Pubblicazione: (2022)
On the Provable Suboptimality of Momentum SGD in Nonstationary Stochastic Optimization
di: Sahu, Sharan, et al.
Pubblicazione: (2026)
di: Sahu, Sharan, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Accelerating Distributed Stochastic Optimization via Self-Repellent Random Walks
di: Hu, Jie, et al.
Pubblicazione: (2024) -
Central Limit Theorem for Two-Timescale Stochastic Approximation with Markovian Noise: Theory and Applications
di: Hu, Jie, et al.
Pubblicazione: (2024) -
Byzantine-Robust Distributed SGD: A Unified Analysis and Tight Error Bounds
di: Ruan, Boyuan, et al.
Pubblicazione: (2026) -
Does SGD really happen in tiny subspaces?
di: Song, Minhak, et al.
Pubblicazione: (2024) -
Demystifying SGD with Doubly Stochastic Gradients
di: Kim, Kyurae, et al.
Pubblicazione: (2024)