Diagonalisation SGD: Fast & Convergent SGD for Non-Differentiable Models via Reparameterisation and Smoothing
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wagner, Dominik, Khajwal, Basim, Ong, C. -H. Luke |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Fast Last-Iterate Convergence of SGD in the Smooth Interpolation Regime
von: Attia, Amit, et al.
Veröffentlicht: (2025)
von: Attia, Amit, et al.
Veröffentlicht: (2025)
Scaling Laws of SignSGD in Linear Regression: When Does It Outperform SGD?
von: Kim, Jihwan, et al.
Veröffentlicht: (2026)
von: Kim, Jihwan, et al.
Veröffentlicht: (2026)
StoSignSGD: Unbiased Structural Stochasticity Fixes SignSGD for Training Large Language Models
von: Yu, Dingzhi, et al.
Veröffentlicht: (2026)
von: Yu, Dingzhi, et al.
Veröffentlicht: (2026)
SGD at the Edge of Stability: The Stochastic Sharpness Gap
von: Liao, Fangshuo, et al.
Veröffentlicht: (2026)
von: Liao, Fangshuo, et al.
Veröffentlicht: (2026)
When and Why SignSGD Outperforms SGD: A Theoretical Study Based on $\ell_1$-norm Lower Bounds
von: Tao, Hongyi, et al.
Veröffentlicht: (2026)
von: Tao, Hongyi, et al.
Veröffentlicht: (2026)
Faster Convergence of Local SGD for Over-Parameterized Models
von: Qin, Tiancheng, et al.
Veröffentlicht: (2022)
von: Qin, Tiancheng, et al.
Veröffentlicht: (2022)
Differentially Private Clipped-SGD: High-Probability Convergence with Arbitrary Clipping Level
von: Khah, Saleh Vatan, et al.
Veröffentlicht: (2025)
von: Khah, Saleh Vatan, et al.
Veröffentlicht: (2025)
From PowerSGD to PowerSGD+: Low-Rank Gradient Compression for Distributed Optimization with Convergence Guarantees
von: Xie, Shengping, et al.
Veröffentlicht: (2025)
von: Xie, Shengping, et al.
Veröffentlicht: (2025)
High-Probability Convergence Guarantees of Decentralized SGD
von: Armacki, Aleksandar, et al.
Veröffentlicht: (2025)
von: Armacki, Aleksandar, et al.
Veröffentlicht: (2025)
Convergence of Clipped-SGD for Convex $(L_0,L_1)$-Smooth Optimization with Heavy-Tailed Noise
von: Chezhegov, Savelii, et al.
Veröffentlicht: (2025)
von: Chezhegov, Savelii, et al.
Veröffentlicht: (2025)
Global Convergence of SGD On Two Layer Neural Nets
von: Gopalani, Pulkit, et al.
Veröffentlicht: (2022)
von: Gopalani, Pulkit, et al.
Veröffentlicht: (2022)
On the Convergence of DP-SGD with Adaptive Clipping
von: Shulgin, Egor, et al.
Veröffentlicht: (2024)
von: Shulgin, Egor, et al.
Veröffentlicht: (2024)
A Hessian-Aware Stochastic Differential Equation for Modelling SGD
von: Li, Xiang, et al.
Veröffentlicht: (2024)
von: Li, Xiang, et al.
Veröffentlicht: (2024)
A Comprehensive Framework for Analyzing the Convergence of Adam: Bridging the Gap with SGD
von: Jin, Ruinan, et al.
Veröffentlicht: (2024)
von: Jin, Ruinan, et al.
Veröffentlicht: (2024)
VAMO: Efficient Zeroth-Order Variance Reduction for SGD with Faster Convergence
von: Chen, Jiahe, et al.
Veröffentlicht: (2025)
von: Chen, Jiahe, et al.
Veröffentlicht: (2025)
Global Convergence of SGD For Logistic Loss on Two Layer Neural Nets
von: Gopalani, Pulkit, et al.
Veröffentlicht: (2023)
von: Gopalani, Pulkit, et al.
Veröffentlicht: (2023)
SLowcal-SGD: Slow Query Points Improve Local-SGD for Stochastic Convex Optimization
von: Dahan, Tehila, et al.
Veröffentlicht: (2023)
von: Dahan, Tehila, et al.
Veröffentlicht: (2023)
Convergence of SGD for Training Neural Networks with Sliced Wasserstein Losses
von: Tanguy, Eloi
Veröffentlicht: (2023)
von: Tanguy, Eloi
Veröffentlicht: (2023)
Sign-SGD via Parameter-Free Optimization
von: Medyakov, Daniil, et al.
Veröffentlicht: (2025)
von: Medyakov, Daniil, et al.
Veröffentlicht: (2025)
Making SGD Parameter-Free
von: Carmon, Yair, et al.
Veröffentlicht: (2022)
von: Carmon, Yair, et al.
Veröffentlicht: (2022)
On the Trajectories of SGD Without Replacement
von: Beneventano, Pierfrancesco
Veröffentlicht: (2023)
von: Beneventano, Pierfrancesco
Veröffentlicht: (2023)
Convergence of SGD with momentum in the nonconvex case: A time window-based analysis
von: Qiu, Junwen, et al.
Veröffentlicht: (2024)
von: Qiu, Junwen, et al.
Veröffentlicht: (2024)
Adaptive SGD with Line-Search and Polyak Stepsizes: Nonconvex Convergence and Accelerated Rates
von: Wu, Haotian
Veröffentlicht: (2025)
von: Wu, Haotian
Veröffentlicht: (2025)
Convergence and concentration properties of constant step-size SGD through Markov chains
von: Merad, Ibrahim, et al.
Veröffentlicht: (2023)
von: Merad, Ibrahim, et al.
Veröffentlicht: (2023)
Tight Long-Term Tail Decay of (Clipped) SGD in Non-Convex Optimization
von: Armacki, Aleksandar, et al.
Veröffentlicht: (2026)
von: Armacki, Aleksandar, et al.
Veröffentlicht: (2026)
Shadowheart SGD: Distributed Asynchronous SGD with Optimal Time Complexity Under Arbitrary Computation and Communication Heterogeneity
von: Tyurin, Alexander, et al.
Veröffentlicht: (2024)
von: Tyurin, Alexander, et al.
Veröffentlicht: (2024)
Demystifying SGD with Doubly Stochastic Gradients
von: Kim, Kyurae, et al.
Veröffentlicht: (2024)
von: Kim, Kyurae, et al.
Veröffentlicht: (2024)
Dimension-adapted Momentum Outscales SGD
von: Ferbach, Damien, et al.
Veröffentlicht: (2025)
von: Ferbach, Damien, et al.
Veröffentlicht: (2025)
Heavy-Tail Phenomenon in Decentralized SGD
von: Gurbuzbalaban, Mert, et al.
Veröffentlicht: (2022)
von: Gurbuzbalaban, Mert, et al.
Veröffentlicht: (2022)
Non-Euclidean SGD for Structured Optimization: Unified Analysis and Improved Rates
von: Kovalev, Dmitry, et al.
Veröffentlicht: (2025)
von: Kovalev, Dmitry, et al.
Veröffentlicht: (2025)
Lower Bounds and Proximally Anchored SGD for Non-Convex Minimization Under Unbounded Variance
von: Fazla, Arda, et al.
Veröffentlicht: (2026)
von: Fazla, Arda, et al.
Veröffentlicht: (2026)
Last-Iterate Convergence of Randomized Kaczmarz and SGD with Greedy Step Size
von: Dereziński, Michał, et al.
Veröffentlicht: (2026)
von: Dereziński, Michał, et al.
Veröffentlicht: (2026)
SGD with memory: fundamental properties and stochastic acceleration
von: Yarotsky, Dmitry, et al.
Veröffentlicht: (2024)
von: Yarotsky, Dmitry, et al.
Veröffentlicht: (2024)
Does SGD really happen in tiny subspaces?
von: Song, Minhak, et al.
Veröffentlicht: (2024)
von: Song, Minhak, et al.
Veröffentlicht: (2024)
The Rich and the Simple: On the Implicit Bias of Adam and SGD
von: Vasudeva, Bhavya, et al.
Veröffentlicht: (2025)
von: Vasudeva, Bhavya, et al.
Veröffentlicht: (2025)
Can SGD Handle Heavy-Tailed Noise?
von: Fatkhullin, Ilyas, et al.
Veröffentlicht: (2025)
von: Fatkhullin, Ilyas, et al.
Veröffentlicht: (2025)
Does SGD Seek Flatness or Sharpness? An Exactly Solvable Model
von: Xu, Yizhou, et al.
Veröffentlicht: (2026)
von: Xu, Yizhou, et al.
Veröffentlicht: (2026)
Non-Smooth Weakly-Convex Finite-sum Coupled Compositional Optimization
von: Hu, Quanqi, et al.
Veröffentlicht: (2023)
von: Hu, Quanqi, et al.
Veröffentlicht: (2023)
SketchySGD: Reliable Stochastic Optimization via Randomized Curvature Estimates
von: Frangella, Zachary, et al.
Veröffentlicht: (2022)
von: Frangella, Zachary, et al.
Veröffentlicht: (2022)
Asynchronous Decentralized SGD under Non-Convexity: A Block-Coordinate Descent Framework
von: Zhou, Yijie, et al.
Veröffentlicht: (2025)
von: Zhou, Yijie, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Fast Last-Iterate Convergence of SGD in the Smooth Interpolation Regime
von: Attia, Amit, et al.
Veröffentlicht: (2025) -
Scaling Laws of SignSGD in Linear Regression: When Does It Outperform SGD?
von: Kim, Jihwan, et al.
Veröffentlicht: (2026) -
StoSignSGD: Unbiased Structural Stochasticity Fixes SignSGD for Training Large Language Models
von: Yu, Dingzhi, et al.
Veröffentlicht: (2026) -
SGD at the Edge of Stability: The Stochastic Sharpness Gap
von: Liao, Fangshuo, et al.
Veröffentlicht: (2026) -
When and Why SignSGD Outperforms SGD: A Theoretical Study Based on $\ell_1$-norm Lower Bounds
von: Tao, Hongyi, et al.
Veröffentlicht: (2026)