Faster Convergence of Local SGD for Over-Parameterized Models
Fuente:
arXiv
Saved in:
| Main Authors: | Qin, Tiancheng, Etesami, S. Rasoul, Uribe, César A. |
|---|---|
| Format: | Preprint |
| Published: |
2022
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Online Reinforcement Learning in Markov Decision Process Using Linear Programming
by: Leon, Vincent, et al.
Published: (2023)
by: Leon, Vincent, et al.
Published: (2023)
VAMO: Efficient Zeroth-Order Variance Reduction for SGD with Faster Convergence
by: Chen, Jiahe, et al.
Published: (2025)
by: Chen, Jiahe, et al.
Published: (2025)
Online Learning for Dynamic Vickrey-Clarke-Groves Mechanism in Unknown Environments
by: Leon, Vincent, et al.
Published: (2025)
by: Leon, Vincent, et al.
Published: (2025)
Toward Global Convergence of Gradient EM for Over-Parameterized Gaussian Mixture Models
by: Xu, Weihang, et al.
Published: (2024)
by: Xu, Weihang, et al.
Published: (2024)
Global Convergence of SGD On Two Layer Neural Nets
by: Gopalani, Pulkit, et al.
Published: (2022)
by: Gopalani, Pulkit, et al.
Published: (2022)
From PowerSGD to PowerSGD+: Low-Rank Gradient Compression for Distributed Optimization with Convergence Guarantees
by: Xie, Shengping, et al.
Published: (2025)
by: Xie, Shengping, et al.
Published: (2025)
Diagonalisation SGD: Fast & Convergent SGD for Non-Differentiable Models via Reparameterisation and Smoothing
by: Wagner, Dominik, et al.
Published: (2024)
by: Wagner, Dominik, et al.
Published: (2024)
Fast Last-Iterate Convergence of SGD in the Smooth Interpolation Regime
by: Attia, Amit, et al.
Published: (2025)
by: Attia, Amit, et al.
Published: (2025)
SLowcal-SGD: Slow Query Points Improve Local-SGD for Stochastic Convex Optimization
by: Dahan, Tehila, et al.
Published: (2023)
by: Dahan, Tehila, et al.
Published: (2023)
A Comprehensive Framework for Analyzing the Convergence of Adam: Bridging the Gap with SGD
by: Jin, Ruinan, et al.
Published: (2024)
by: Jin, Ruinan, et al.
Published: (2024)
Global Convergence of SGD For Logistic Loss on Two Layer Neural Nets
by: Gopalani, Pulkit, et al.
Published: (2023)
by: Gopalani, Pulkit, et al.
Published: (2023)
On the Convergence of DP-SGD with Adaptive Clipping
by: Shulgin, Egor, et al.
Published: (2024)
by: Shulgin, Egor, et al.
Published: (2024)
Drop-Muon: Update Less, Converge Faster
by: Gruntkowska, Kaja, et al.
Published: (2025)
by: Gruntkowska, Kaja, et al.
Published: (2025)
Convergence of SGD with momentum in the nonconvex case: A time window-based analysis
by: Qiu, Junwen, et al.
Published: (2024)
by: Qiu, Junwen, et al.
Published: (2024)
Differentially Private Clipped-SGD: High-Probability Convergence with Arbitrary Clipping Level
by: Khah, Saleh Vatan, et al.
Published: (2025)
by: Khah, Saleh Vatan, et al.
Published: (2025)
Adaptive SGD with Line-Search and Polyak Stepsizes: Nonconvex Convergence and Accelerated Rates
by: Wu, Haotian
Published: (2025)
by: Wu, Haotian
Published: (2025)
Convergence and concentration properties of constant step-size SGD through Markov chains
by: Merad, Ibrahim, et al.
Published: (2023)
by: Merad, Ibrahim, et al.
Published: (2023)
Distributed Nash Equilibrium Seeking in Non-Monotone Games over the Simplex
by: Tatarenko, Tatiana, et al.
Published: (2025)
by: Tatarenko, Tatiana, et al.
Published: (2025)
Convergence of SGD for Training Neural Networks with Sliced Wasserstein Losses
by: Tanguy, Eloi
Published: (2023)
by: Tanguy, Eloi
Published: (2023)
High-Probability Convergence Guarantees of Decentralized SGD
by: Armacki, Aleksandar, et al.
Published: (2025)
by: Armacki, Aleksandar, et al.
Published: (2025)
A Fixed Point Framework for the Existence of EFX Allocations
by: Etesami, S. Rasoul
Published: (2025)
by: Etesami, S. Rasoul
Published: (2025)
Preconditioned Gradient Descent for Over-Parameterized Nonconvex Matrix Factorization
by: Zhang, Gavin, et al.
Published: (2025)
by: Zhang, Gavin, et al.
Published: (2025)
Provably Convergent Federated Trilevel Learning
by: Jiao, Yang, et al.
Published: (2023)
by: Jiao, Yang, et al.
Published: (2023)
Faster Convergence of Stochastic Accelerated Gradient Descent under Interpolation
by: Mishkin, Aaron, et al.
Published: (2024)
by: Mishkin, Aaron, et al.
Published: (2024)
Incremental Quasi-Newton Methods with Faster Superlinear Convergence Rates
by: Liu, Zhuanghua, et al.
Published: (2024)
by: Liu, Zhuanghua, et al.
Published: (2024)
Learning How to Strategically Disclose Information
by: Velicheti, Raj Kiriti, et al.
Published: (2024)
by: Velicheti, Raj Kiriti, et al.
Published: (2024)
Understanding Outer Optimizers in Local SGD: Learning Rates, Momentum, and Acceleration
by: Khaled, Ahmed, et al.
Published: (2025)
by: Khaled, Ahmed, et al.
Published: (2025)
Faster Convergence of Riemannian Stochastic Gradient Descent with Increasing Batch Size
by: Oowada, Kanata, et al.
Published: (2025)
by: Oowada, Kanata, et al.
Published: (2025)
FastPart: Over-Parameterized Stochastic Gradient Descent for Sparse optimisation on Measures
by: De Castro, Yohann, et al.
Published: (2023)
by: De Castro, Yohann, et al.
Published: (2023)
Convergence of Clipped-SGD for Convex $(L_0,L_1)$-Smooth Optimization with Heavy-Tailed Noise
by: Chezhegov, Savelii, et al.
Published: (2025)
by: Chezhegov, Savelii, et al.
Published: (2025)
Learning Over-Relaxation Policies for ADMM with Convergence Guarantees
by: Lin, Junan, et al.
Published: (2026)
by: Lin, Junan, et al.
Published: (2026)
Last-Iterate Convergence of Randomized Kaczmarz and SGD with Greedy Step Size
by: Dereziński, Michał, et al.
Published: (2026)
by: Dereziński, Michał, et al.
Published: (2026)
Making SGD Parameter-Free
by: Carmon, Yair, et al.
Published: (2022)
by: Carmon, Yair, et al.
Published: (2022)
On the Trajectories of SGD Without Replacement
by: Beneventano, Pierfrancesco
Published: (2023)
by: Beneventano, Pierfrancesco
Published: (2023)
Attention-Enhanced Graph Filtering for False Data Injection Attack Detection and Localization
by: Abdulin, Ruslan, et al.
Published: (2026)
by: Abdulin, Ruslan, et al.
Published: (2026)
A Hessian-Aware Stochastic Differential Equation for Modelling SGD
by: Li, Xiang, et al.
Published: (2024)
by: Li, Xiang, et al.
Published: (2024)
Does SGD Seek Flatness or Sharpness? An Exactly Solvable Model
by: Xu, Yizhou, et al.
Published: (2026)
by: Xu, Yizhou, et al.
Published: (2026)
Shadowheart SGD: Distributed Asynchronous SGD with Optimal Time Complexity Under Arbitrary Computation and Communication Heterogeneity
by: Tyurin, Alexander, et al.
Published: (2024)
by: Tyurin, Alexander, et al.
Published: (2024)
Heavy-Tail Phenomenon in Decentralized SGD
by: Gurbuzbalaban, Mert, et al.
Published: (2022)
by: Gurbuzbalaban, Mert, et al.
Published: (2022)
Dimension-adapted Momentum Outscales SGD
by: Ferbach, Damien, et al.
Published: (2025)
by: Ferbach, Damien, et al.
Published: (2025)
Similar Items
-
Online Reinforcement Learning in Markov Decision Process Using Linear Programming
by: Leon, Vincent, et al.
Published: (2023) -
VAMO: Efficient Zeroth-Order Variance Reduction for SGD with Faster Convergence
by: Chen, Jiahe, et al.
Published: (2025) -
Online Learning for Dynamic Vickrey-Clarke-Groves Mechanism in Unknown Environments
by: Leon, Vincent, et al.
Published: (2025) -
Toward Global Convergence of Gradient EM for Over-Parameterized Gaussian Mixture Models
by: Xu, Weihang, et al.
Published: (2024) -
Global Convergence of SGD On Two Layer Neural Nets
by: Gopalani, Pulkit, et al.
Published: (2022)