Gespeichert in:
| Hauptverfasser: | Chen, Haoran, Wang, Wentao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2604.25550 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SignSGD with Federated Voting
von: Park, Chanho, et al.
Veröffentlicht: (2024)
von: Park, Chanho, et al.
Veröffentlicht: (2024)
Phases of Muon: When Muon Eclipses SignSGD
von: Paquette, Elliot, et al.
Veröffentlicht: (2026)
von: Paquette, Elliot, et al.
Veröffentlicht: (2026)
Scaling Laws of SignSGD in Linear Regression: When Does It Outperform SGD?
von: Kim, Jihwan, et al.
Veröffentlicht: (2026)
von: Kim, Jihwan, et al.
Veröffentlicht: (2026)
StoSignSGD: Unbiased Structural Stochasticity Fixes SignSGD for Training Large Language Models
von: Yu, Dingzhi, et al.
Veröffentlicht: (2026)
von: Yu, Dingzhi, et al.
Veröffentlicht: (2026)
SignSGD with Federated Defense: Harnessing Adversarial Attacks through Gradient Sign Decoding
von: Park, Chanho, et al.
Veröffentlicht: (2024)
von: Park, Chanho, et al.
Veröffentlicht: (2024)
Leveraging Coordinate Momentum in SignSGD and Muon: Memory-Optimized Zero-Order
von: Petrov, Egor, et al.
Veröffentlicht: (2025)
von: Petrov, Egor, et al.
Veröffentlicht: (2025)
Hierarchical Federated Learning with SignSGD: A Highly Communication-Efficient Approach
von: Kazemi, Amirreza, et al.
Veröffentlicht: (2026)
von: Kazemi, Amirreza, et al.
Veröffentlicht: (2026)
When and Why SignSGD Outperforms SGD: A Theoretical Study Based on $\ell_1$-norm Lower Bounds
von: Tao, Hongyi, et al.
Veröffentlicht: (2026)
von: Tao, Hongyi, et al.
Veröffentlicht: (2026)
Convergence Analysis of SGD under Expected Smoothness
von: Kawamoto, Yuta, et al.
Veröffentlicht: (2025)
von: Kawamoto, Yuta, et al.
Veröffentlicht: (2025)
Small Batch Size Training for Language Models: When Vanilla SGD Works, and Why Gradient Accumulation Is Wasteful
von: Marek, Martin, et al.
Veröffentlicht: (2025)
von: Marek, Martin, et al.
Veröffentlicht: (2025)
PCDP-SGD: Improving the Convergence of Differentially Private SGD via Projection in Advance
von: Sha, Haichao, et al.
Veröffentlicht: (2023)
von: Sha, Haichao, et al.
Veröffentlicht: (2023)
Perfect Parallelization in Mini-Batch SGD with Classical Momentum Acceleration
von: Garg, Sachin, et al.
Veröffentlicht: (2026)
von: Garg, Sachin, et al.
Veröffentlicht: (2026)
Stochastic-Sign SGD for Federated Learning with Theoretical Guarantees
von: Jin, Richeng, et al.
Veröffentlicht: (2020)
von: Jin, Richeng, et al.
Veröffentlicht: (2020)
Sign-SGD via Parameter-Free Optimization
von: Medyakov, Daniil, et al.
Veröffentlicht: (2025)
von: Medyakov, Daniil, et al.
Veröffentlicht: (2025)
Demystifying Why Local Aggregation Helps: Convergence Analysis of Hierarchical SGD
von: Wang, Jiayi, et al.
Veröffentlicht: (2020)
von: Wang, Jiayi, et al.
Veröffentlicht: (2020)
From PowerSGD to PowerSGD+: Low-Rank Gradient Compression for Distributed Optimization with Convergence Guarantees
von: Xie, Shengping, et al.
Veröffentlicht: (2025)
von: Xie, Shengping, et al.
Veröffentlicht: (2025)
From One-Pass SGD to Data Reuse: Mini-Batch Scaling Laws in Sketched Linear Regression
von: Chen, Ziyan, et al.
Veröffentlicht: (2026)
von: Chen, Ziyan, et al.
Veröffentlicht: (2026)
On the Convergence of DP-SGD with Adaptive Clipping
von: Shulgin, Egor, et al.
Veröffentlicht: (2024)
von: Shulgin, Egor, et al.
Veröffentlicht: (2024)
Error estimates between SGD with momentum and underdamped Langevin diffusion
von: Guillin, Arnaud, et al.
Veröffentlicht: (2024)
von: Guillin, Arnaud, et al.
Veröffentlicht: (2024)
SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam
von: Peng, Hanyang, et al.
Veröffentlicht: (2025)
von: Peng, Hanyang, et al.
Veröffentlicht: (2025)
The Marginal Value of Momentum for Small Learning Rate SGD
von: Wang, Runzhe, et al.
Veröffentlicht: (2023)
von: Wang, Runzhe, et al.
Veröffentlicht: (2023)
VAMO: Efficient Zeroth-Order Variance Reduction for SGD with Faster Convergence
von: Chen, Jiahe, et al.
Veröffentlicht: (2025)
von: Chen, Jiahe, et al.
Veröffentlicht: (2025)
SGD for Variational Inference: Tackling Unbounded Variance via Preconditioning and Dynamic Batching
von: Labarrière, Hippolyte, et al.
Veröffentlicht: (2026)
von: Labarrière, Hippolyte, et al.
Veröffentlicht: (2026)
Diagonalisation SGD: Fast & Convergent SGD for Non-Differentiable Models via Reparameterisation and Smoothing
von: Wagner, Dominik, et al.
Veröffentlicht: (2024)
von: Wagner, Dominik, et al.
Veröffentlicht: (2024)
Faster Convergence of Local SGD for Over-Parameterized Models
von: Qin, Tiancheng, et al.
Veröffentlicht: (2022)
von: Qin, Tiancheng, et al.
Veröffentlicht: (2022)
Global Convergence of SGD On Two Layer Neural Nets
von: Gopalani, Pulkit, et al.
Veröffentlicht: (2022)
von: Gopalani, Pulkit, et al.
Veröffentlicht: (2022)
High-Probability Convergence Guarantees of Decentralized SGD
von: Armacki, Aleksandar, et al.
Veröffentlicht: (2025)
von: Armacki, Aleksandar, et al.
Veröffentlicht: (2025)
A Comprehensive Framework for Analyzing the Convergence of Adam: Bridging the Gap with SGD
von: Jin, Ruinan, et al.
Veröffentlicht: (2024)
von: Jin, Ruinan, et al.
Veröffentlicht: (2024)
Optimal Growth Schedules for Batch Size and Learning Rate in SGD that Reduce SFO Complexity
von: Umeda, Hikaru, et al.
Veröffentlicht: (2025)
von: Umeda, Hikaru, et al.
Veröffentlicht: (2025)
Switching the Loss Reduces the Cost in Batch (Offline) Reinforcement Learning
von: Ayoub, Alex, et al.
Veröffentlicht: (2024)
von: Ayoub, Alex, et al.
Veröffentlicht: (2024)
GT-Space: Enhancing Heterogeneous Collaborative Perception with Ground Truth Feature Space
von: Wang, Wentao, et al.
Veröffentlicht: (2026)
von: Wang, Wentao, et al.
Veröffentlicht: (2026)
Fast Last-Iterate Convergence of SGD in the Smooth Interpolation Regime
von: Attia, Amit, et al.
Veröffentlicht: (2025)
von: Attia, Amit, et al.
Veröffentlicht: (2025)
Convergent Privacy Loss of Noisy-SGD without Convexity and Smoothness
von: Chien, Eli, et al.
Veröffentlicht: (2024)
von: Chien, Eli, et al.
Veröffentlicht: (2024)
Online Linear Programming with Batching
von: Xu, Haoran, et al.
Veröffentlicht: (2024)
von: Xu, Haoran, et al.
Veröffentlicht: (2024)
Convergence, Sticking and Escape: Stochastic Dynamics Near Critical Points in SGD
von: Dudukalov, Dmitry, et al.
Veröffentlicht: (2025)
von: Dudukalov, Dmitry, et al.
Veröffentlicht: (2025)
Hybrid Unsupervised Learning Strategy for Monitoring Industrial Batch Processes
von: Frey, Christian W.
Veröffentlicht: (2024)
von: Frey, Christian W.
Veröffentlicht: (2024)
Full-Batch Gradient Descent Outperforms One-Pass SGD: Sample Complexity Separation in Single-Index Learning
von: Kovačević, Filip, et al.
Veröffentlicht: (2026)
von: Kovačević, Filip, et al.
Veröffentlicht: (2026)
Noise is All You Need: Private Second-Order Convergence of Noisy SGD
von: Avdiukhin, Dmitrii, et al.
Veröffentlicht: (2024)
von: Avdiukhin, Dmitrii, et al.
Veröffentlicht: (2024)
Convergence Bound and Critical Batch Size of Muon Optimizer
von: Sato, Naoki, et al.
Veröffentlicht: (2025)
von: Sato, Naoki, et al.
Veröffentlicht: (2025)
Convergence of SGD for Training Neural Networks with Sliced Wasserstein Losses
von: Tanguy, Eloi
Veröffentlicht: (2023)
von: Tanguy, Eloi
Veröffentlicht: (2023)
Ähnliche Einträge
-
SignSGD with Federated Voting
von: Park, Chanho, et al.
Veröffentlicht: (2024) -
Phases of Muon: When Muon Eclipses SignSGD
von: Paquette, Elliot, et al.
Veröffentlicht: (2026) -
Scaling Laws of SignSGD in Linear Regression: When Does It Outperform SGD?
von: Kim, Jihwan, et al.
Veröffentlicht: (2026) -
StoSignSGD: Unbiased Structural Stochasticity Fixes SignSGD for Training Large Language Models
von: Yu, Dingzhi, et al.
Veröffentlicht: (2026) -
SignSGD with Federated Defense: Harnessing Adversarial Attacks through Gradient Sign Decoding
von: Park, Chanho, et al.
Veröffentlicht: (2024)