A Precise Characterization of SGD Stability Using Loss Surface Geometry
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Dexter, Gregory, Ocejo, Borja, Keerthi, Sathiya, Gupta, Aman, Acharya, Ayan, Khanna, Rajiv |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Edge of Stochastic Stability: Revisiting the Edge of Stability for SGD
von: Andreyev, Arseniy, et al.
Veröffentlicht: (2024)
von: Andreyev, Arseniy, et al.
Veröffentlicht: (2024)
optimizn: a Python Library for Developing Customized Optimization Algorithms
von: Sathiya, Akshay, et al.
Veröffentlicht: (2025)
von: Sathiya, Akshay, et al.
Veröffentlicht: (2025)
Global Convergence of SGD For Logistic Loss on Two Layer Neural Nets
von: Gopalani, Pulkit, et al.
Veröffentlicht: (2023)
von: Gopalani, Pulkit, et al.
Veröffentlicht: (2023)
Convergence of SGD for Training Neural Networks with Sliced Wasserstein Losses
von: Tanguy, Eloi
Veröffentlicht: (2023)
von: Tanguy, Eloi
Veröffentlicht: (2023)
Perturbed Iterate SGD for Lipschitz Continuous Loss Functions with Numerical Error and Adaptive Step Sizes
von: Metel, Michael R.
Veröffentlicht: (2022)
von: Metel, Michael R.
Veröffentlicht: (2022)
SGD at the Edge of Stability: The Stochastic Sharpness Gap
von: Liao, Fangshuo, et al.
Veröffentlicht: (2026)
von: Liao, Fangshuo, et al.
Veröffentlicht: (2026)
SLowcal-SGD: Slow Query Points Improve Local-SGD for Stochastic Convex Optimization
von: Dahan, Tehila, et al.
Veröffentlicht: (2023)
von: Dahan, Tehila, et al.
Veröffentlicht: (2023)
Making SGD Parameter-Free
von: Carmon, Yair, et al.
Veröffentlicht: (2022)
von: Carmon, Yair, et al.
Veröffentlicht: (2022)
On the Trajectories of SGD Without Replacement
von: Beneventano, Pierfrancesco
Veröffentlicht: (2023)
von: Beneventano, Pierfrancesco
Veröffentlicht: (2023)
Enhancing Stochastic Optimization for Statistical Efficiency Using ROOT-SGD with Diminishing Stepsize
von: Li, Chris Junchi
Veröffentlicht: (2024)
von: Li, Chris Junchi
Veröffentlicht: (2024)
From PowerSGD to PowerSGD+: Low-Rank Gradient Compression for Distributed Optimization with Convergence Guarantees
von: Xie, Shengping, et al.
Veröffentlicht: (2025)
von: Xie, Shengping, et al.
Veröffentlicht: (2025)
Shadowheart SGD: Distributed Asynchronous SGD with Optimal Time Complexity Under Arbitrary Computation and Communication Heterogeneity
von: Tyurin, Alexander, et al.
Veröffentlicht: (2024)
von: Tyurin, Alexander, et al.
Veröffentlicht: (2024)
Dimension-adapted Momentum Outscales SGD
von: Ferbach, Damien, et al.
Veröffentlicht: (2025)
von: Ferbach, Damien, et al.
Veröffentlicht: (2025)
Heavy-Tail Phenomenon in Decentralized SGD
von: Gurbuzbalaban, Mert, et al.
Veröffentlicht: (2022)
von: Gurbuzbalaban, Mert, et al.
Veröffentlicht: (2022)
Demystifying SGD with Doubly Stochastic Gradients
von: Kim, Kyurae, et al.
Veröffentlicht: (2024)
von: Kim, Kyurae, et al.
Veröffentlicht: (2024)
The Rich and the Simple: On the Implicit Bias of Adam and SGD
von: Vasudeva, Bhavya, et al.
Veröffentlicht: (2025)
von: Vasudeva, Bhavya, et al.
Veröffentlicht: (2025)
Sign-SGD via Parameter-Free Optimization
von: Medyakov, Daniil, et al.
Veröffentlicht: (2025)
von: Medyakov, Daniil, et al.
Veröffentlicht: (2025)
Can SGD Handle Heavy-Tailed Noise?
von: Fatkhullin, Ilyas, et al.
Veröffentlicht: (2025)
von: Fatkhullin, Ilyas, et al.
Veröffentlicht: (2025)
SGD with memory: fundamental properties and stochastic acceleration
von: Yarotsky, Dmitry, et al.
Veröffentlicht: (2024)
von: Yarotsky, Dmitry, et al.
Veröffentlicht: (2024)
Does SGD really happen in tiny subspaces?
von: Song, Minhak, et al.
Veröffentlicht: (2024)
von: Song, Minhak, et al.
Veröffentlicht: (2024)
A Simplified Analysis of SGD for Linear Regression with Weight Averaging
von: Meterez, Alexandru, et al.
Veröffentlicht: (2025)
von: Meterez, Alexandru, et al.
Veröffentlicht: (2025)
A Hessian-Aware Stochastic Differential Equation for Modelling SGD
von: Li, Xiang, et al.
Veröffentlicht: (2024)
von: Li, Xiang, et al.
Veröffentlicht: (2024)
From Gradient Clipping to Normalization for Heavy Tailed SGD
von: Hübler, Florian, et al.
Veröffentlicht: (2024)
von: Hübler, Florian, et al.
Veröffentlicht: (2024)
The Optimality of (Accelerated) SGD for High-Dimensional Quadratic Optimization
von: Zhang, Haihan, et al.
Veröffentlicht: (2024)
von: Zhang, Haihan, et al.
Veröffentlicht: (2024)
Dual-Delayed Asynchronous SGD for Arbitrarily Heterogeneous Data
von: Wang, Xiaolu, et al.
Veröffentlicht: (2024)
von: Wang, Xiaolu, et al.
Veröffentlicht: (2024)
Faster Convergence of Local SGD for Over-Parameterized Models
von: Qin, Tiancheng, et al.
Veröffentlicht: (2022)
von: Qin, Tiancheng, et al.
Veröffentlicht: (2022)
The Marginal Value of Momentum for Small Learning Rate SGD
von: Wang, Runzhe, et al.
Veröffentlicht: (2023)
von: Wang, Runzhe, et al.
Veröffentlicht: (2023)
Phases of Muon: When Muon Eclipses SignSGD
von: Paquette, Elliot, et al.
Veröffentlicht: (2026)
von: Paquette, Elliot, et al.
Veröffentlicht: (2026)
SGD with Partial Hessian for Deep Neural Networks Optimization
von: Sun, Ying, et al.
Veröffentlicht: (2024)
von: Sun, Ying, et al.
Veröffentlicht: (2024)
Optimal Projection-Free Adaptive SGD for Matrix Optimization
von: Kovalev, Dmitry
Veröffentlicht: (2026)
von: Kovalev, Dmitry
Veröffentlicht: (2026)
SGD with Adaptive Preconditioning: Unified Analysis and Momentum Acceleration
von: Kovalev, Dmitry
Veröffentlicht: (2025)
von: Kovalev, Dmitry
Veröffentlicht: (2025)
Global Convergence of SGD On Two Layer Neural Nets
von: Gopalani, Pulkit, et al.
Veröffentlicht: (2022)
von: Gopalani, Pulkit, et al.
Veröffentlicht: (2022)
On the Provable Suboptimality of Momentum SGD in Nonstationary Stochastic Optimization
von: Sahu, Sharan, et al.
Veröffentlicht: (2026)
von: Sahu, Sharan, et al.
Veröffentlicht: (2026)
Accelerating Single-Pass SGD for Generalized Linear Prediction
von: Chen, Qian, et al.
Veröffentlicht: (2026)
von: Chen, Qian, et al.
Veröffentlicht: (2026)
A Comprehensive Framework for Analyzing the Convergence of Adam: Bridging the Gap with SGD
von: Jin, Ruinan, et al.
Veröffentlicht: (2024)
von: Jin, Ruinan, et al.
Veröffentlicht: (2024)
Bias-Optimal Bounds for SGD: A Computer-Aided Lyapunov Analysis
von: Cortild, Daniel, et al.
Veröffentlicht: (2025)
von: Cortild, Daniel, et al.
Veröffentlicht: (2025)
Proactive DP: A Multple Target Optimization Framework for DP-SGD
von: van Dijk, Marten, et al.
Veröffentlicht: (2021)
von: van Dijk, Marten, et al.
Veröffentlicht: (2021)
Fast Last-Iterate Convergence of SGD in the Smooth Interpolation Regime
von: Attia, Amit, et al.
Veröffentlicht: (2025)
von: Attia, Amit, et al.
Veröffentlicht: (2025)
Does SGD Seek Flatness or Sharpness? An Exactly Solvable Model
von: Xu, Yizhou, et al.
Veröffentlicht: (2026)
von: Xu, Yizhou, et al.
Veröffentlicht: (2026)
Byzantine-Robust Distributed SGD: A Unified Analysis and Tight Error Bounds
von: Ruan, Boyuan, et al.
Veröffentlicht: (2026)
von: Ruan, Boyuan, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Edge of Stochastic Stability: Revisiting the Edge of Stability for SGD
von: Andreyev, Arseniy, et al.
Veröffentlicht: (2024) -
optimizn: a Python Library for Developing Customized Optimization Algorithms
von: Sathiya, Akshay, et al.
Veröffentlicht: (2025) -
Global Convergence of SGD For Logistic Loss on Two Layer Neural Nets
von: Gopalani, Pulkit, et al.
Veröffentlicht: (2023) -
Convergence of SGD for Training Neural Networks with Sliced Wasserstein Losses
von: Tanguy, Eloi
Veröffentlicht: (2023) -
Perturbed Iterate SGD for Lipschitz Continuous Loss Functions with Numerical Error and Adaptive Step Sizes
von: Metel, Michael R.
Veröffentlicht: (2022)