Enhancing SignSGD: Small-Batch Convergence Analysis and a Hybrid Switching Strategy
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Chen, Haoran, Wang, Wentao |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
SignSGD with Federated Voting
par: Park, Chanho, et autres
Publié: (2024)
par: Park, Chanho, et autres
Publié: (2024)
Phases of Muon: When Muon Eclipses SignSGD
par: Paquette, Elliot, et autres
Publié: (2026)
par: Paquette, Elliot, et autres
Publié: (2026)
Scaling Laws of SignSGD in Linear Regression: When Does It Outperform SGD?
par: Kim, Jihwan, et autres
Publié: (2026)
par: Kim, Jihwan, et autres
Publié: (2026)
StoSignSGD: Unbiased Structural Stochasticity Fixes SignSGD for Training Large Language Models
par: Yu, Dingzhi, et autres
Publié: (2026)
par: Yu, Dingzhi, et autres
Publié: (2026)
SignSGD with Federated Defense: Harnessing Adversarial Attacks through Gradient Sign Decoding
par: Park, Chanho, et autres
Publié: (2024)
par: Park, Chanho, et autres
Publié: (2024)
Leveraging Coordinate Momentum in SignSGD and Muon: Memory-Optimized Zero-Order
par: Petrov, Egor, et autres
Publié: (2025)
par: Petrov, Egor, et autres
Publié: (2025)
Hierarchical Federated Learning with SignSGD: A Highly Communication-Efficient Approach
par: Kazemi, Amirreza, et autres
Publié: (2026)
par: Kazemi, Amirreza, et autres
Publié: (2026)
When and Why SignSGD Outperforms SGD: A Theoretical Study Based on $\ell_1$-norm Lower Bounds
par: Tao, Hongyi, et autres
Publié: (2026)
par: Tao, Hongyi, et autres
Publié: (2026)
Convergence Analysis of SGD under Expected Smoothness
par: Kawamoto, Yuta, et autres
Publié: (2025)
par: Kawamoto, Yuta, et autres
Publié: (2025)
Small Batch Size Training for Language Models: When Vanilla SGD Works, and Why Gradient Accumulation Is Wasteful
par: Marek, Martin, et autres
Publié: (2025)
par: Marek, Martin, et autres
Publié: (2025)
Perfect Parallelization in Mini-Batch SGD with Classical Momentum Acceleration
par: Garg, Sachin, et autres
Publié: (2026)
par: Garg, Sachin, et autres
Publié: (2026)
PCDP-SGD: Improving the Convergence of Differentially Private SGD via Projection in Advance
par: Sha, Haichao, et autres
Publié: (2023)
par: Sha, Haichao, et autres
Publié: (2023)
Stochastic-Sign SGD for Federated Learning with Theoretical Guarantees
par: Jin, Richeng, et autres
Publié: (2020)
par: Jin, Richeng, et autres
Publié: (2020)
Sign-SGD via Parameter-Free Optimization
par: Medyakov, Daniil, et autres
Publié: (2025)
par: Medyakov, Daniil, et autres
Publié: (2025)
From PowerSGD to PowerSGD+: Low-Rank Gradient Compression for Distributed Optimization with Convergence Guarantees
par: Xie, Shengping, et autres
Publié: (2025)
par: Xie, Shengping, et autres
Publié: (2025)
From One-Pass SGD to Data Reuse: Mini-Batch Scaling Laws in Sketched Linear Regression
par: Chen, Ziyan, et autres
Publié: (2026)
par: Chen, Ziyan, et autres
Publié: (2026)
Demystifying Why Local Aggregation Helps: Convergence Analysis of Hierarchical SGD
par: Wang, Jiayi, et autres
Publié: (2020)
par: Wang, Jiayi, et autres
Publié: (2020)
SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam
par: Peng, Hanyang, et autres
Publié: (2025)
par: Peng, Hanyang, et autres
Publié: (2025)
On the Convergence of DP-SGD with Adaptive Clipping
par: Shulgin, Egor, et autres
Publié: (2024)
par: Shulgin, Egor, et autres
Publié: (2024)
VAMO: Efficient Zeroth-Order Variance Reduction for SGD with Faster Convergence
par: Chen, Jiahe, et autres
Publié: (2025)
par: Chen, Jiahe, et autres
Publié: (2025)
The Marginal Value of Momentum for Small Learning Rate SGD
par: Wang, Runzhe, et autres
Publié: (2023)
par: Wang, Runzhe, et autres
Publié: (2023)
Error estimates between SGD with momentum and underdamped Langevin diffusion
par: Guillin, Arnaud, et autres
Publié: (2024)
par: Guillin, Arnaud, et autres
Publié: (2024)
SGD for Variational Inference: Tackling Unbounded Variance via Preconditioning and Dynamic Batching
par: Labarrière, Hippolyte, et autres
Publié: (2026)
par: Labarrière, Hippolyte, et autres
Publié: (2026)
Faster Convergence of Local SGD for Over-Parameterized Models
par: Qin, Tiancheng, et autres
Publié: (2022)
par: Qin, Tiancheng, et autres
Publié: (2022)
Global Convergence of SGD On Two Layer Neural Nets
par: Gopalani, Pulkit, et autres
Publié: (2022)
par: Gopalani, Pulkit, et autres
Publié: (2022)
Switching the Loss Reduces the Cost in Batch (Offline) Reinforcement Learning
par: Ayoub, Alex, et autres
Publié: (2024)
par: Ayoub, Alex, et autres
Publié: (2024)
Diagonalisation SGD: Fast & Convergent SGD for Non-Differentiable Models via Reparameterisation and Smoothing
par: Wagner, Dominik, et autres
Publié: (2024)
par: Wagner, Dominik, et autres
Publié: (2024)
A Comprehensive Framework for Analyzing the Convergence of Adam: Bridging the Gap with SGD
par: Jin, Ruinan, et autres
Publié: (2024)
par: Jin, Ruinan, et autres
Publié: (2024)
Optimal Growth Schedules for Batch Size and Learning Rate in SGD that Reduce SFO Complexity
par: Umeda, Hikaru, et autres
Publié: (2025)
par: Umeda, Hikaru, et autres
Publié: (2025)
High-Probability Convergence Guarantees of Decentralized SGD
par: Armacki, Aleksandar, et autres
Publié: (2025)
par: Armacki, Aleksandar, et autres
Publié: (2025)
Full-Batch Gradient Descent Outperforms One-Pass SGD: Sample Complexity Separation in Single-Index Learning
par: Kovačević, Filip, et autres
Publié: (2026)
par: Kovačević, Filip, et autres
Publié: (2026)
Fast Last-Iterate Convergence of SGD in the Smooth Interpolation Regime
par: Attia, Amit, et autres
Publié: (2025)
par: Attia, Amit, et autres
Publié: (2025)
Convergent Privacy Loss of Noisy-SGD without Convexity and Smoothness
par: Chien, Eli, et autres
Publié: (2024)
par: Chien, Eli, et autres
Publié: (2024)
Convergence Bound and Critical Batch Size of Muon Optimizer
par: Sato, Naoki, et autres
Publié: (2025)
par: Sato, Naoki, et autres
Publié: (2025)
Noise is All You Need: Private Second-Order Convergence of Noisy SGD
par: Avdiukhin, Dmitrii, et autres
Publié: (2024)
par: Avdiukhin, Dmitrii, et autres
Publié: (2024)
Convergence, Sticking and Escape: Stochastic Dynamics Near Critical Points in SGD
par: Dudukalov, Dmitry, et autres
Publié: (2025)
par: Dudukalov, Dmitry, et autres
Publié: (2025)
Hybrid Unsupervised Learning Strategy for Monitoring Industrial Batch Processes
par: Frey, Christian W.
Publié: (2024)
par: Frey, Christian W.
Publié: (2024)
Online Linear Programming with Batching
par: Xu, Haoran, et autres
Publié: (2024)
par: Xu, Haoran, et autres
Publié: (2024)
GT-Space: Enhancing Heterogeneous Collaborative Perception with Ground Truth Feature Space
par: Wang, Wentao, et autres
Publié: (2026)
par: Wang, Wentao, et autres
Publié: (2026)
Convergence of SGD for Training Neural Networks with Sliced Wasserstein Losses
par: Tanguy, Eloi
Publié: (2023)
par: Tanguy, Eloi
Publié: (2023)
Documents similaires
-
SignSGD with Federated Voting
par: Park, Chanho, et autres
Publié: (2024) -
Phases of Muon: When Muon Eclipses SignSGD
par: Paquette, Elliot, et autres
Publié: (2026) -
Scaling Laws of SignSGD in Linear Regression: When Does It Outperform SGD?
par: Kim, Jihwan, et autres
Publié: (2026) -
StoSignSGD: Unbiased Structural Stochasticity Fixes SignSGD for Training Large Language Models
par: Yu, Dingzhi, et autres
Publié: (2026) -
SignSGD with Federated Defense: Harnessing Adversarial Attacks through Gradient Sign Decoding
par: Park, Chanho, et autres
Publié: (2024)