Edge of Stochastic Stability: Revisiting the Edge of Stability for SGD
Fuente:
arXiv
Salvato in:
| Autori principali: | Andreyev, Arseniy, Beneventano, Pierfrancesco |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Momentum Further Constrains Sharpness at the Edge of Stochastic Stability
di: Andreyev, Arseniy, et al.
Pubblicazione: (2026)
di: Andreyev, Arseniy, et al.
Pubblicazione: (2026)
On the Trajectories of SGD Without Replacement
di: Beneventano, Pierfrancesco
Pubblicazione: (2023)
di: Beneventano, Pierfrancesco
Pubblicazione: (2023)
How Neural Networks Learn the Support is an Implicit Regularization Effect of SGD
di: Beneventano, Pierfrancesco, et al.
Pubblicazione: (2024)
di: Beneventano, Pierfrancesco, et al.
Pubblicazione: (2024)
SGD at the Edge of Stability: The Stochastic Sharpness Gap
di: Liao, Fangshuo, et al.
Pubblicazione: (2026)
di: Liao, Fangshuo, et al.
Pubblicazione: (2026)
Does SGD Seek Flatness or Sharpness? An Exactly Solvable Model
di: Xu, Yizhou, et al.
Pubblicazione: (2026)
di: Xu, Yizhou, et al.
Pubblicazione: (2026)
Does Weight Decay Enhance Training Stability?
di: Saether, Marius, et al.
Pubblicazione: (2026)
di: Saether, Marius, et al.
Pubblicazione: (2026)
Gradient Descent Converges Linearly to Flatter Minima than Gradient Flow in Shallow Linear Networks
di: Beneventano, Pierfrancesco, et al.
Pubblicazione: (2025)
di: Beneventano, Pierfrancesco, et al.
Pubblicazione: (2025)
Too Sharp, Too Sure: When Calibration Follows Curvature
di: Morosini, Alessandro, et al.
Pubblicazione: (2026)
di: Morosini, Alessandro, et al.
Pubblicazione: (2026)
Non-Euclidean Gradient Descent Operates at the Edge of Stability
di: Islamov, Rustem, et al.
Pubblicazione: (2026)
di: Islamov, Rustem, et al.
Pubblicazione: (2026)
Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias
di: Das, Mohua, et al.
Pubblicazione: (2026)
di: Das, Mohua, et al.
Pubblicazione: (2026)
Zeroth-Order Optimization at the Edge of Stability
di: Song, Minhak, et al.
Pubblicazione: (2026)
di: Song, Minhak, et al.
Pubblicazione: (2026)
Criteria and Bias of Parameterized Linear Regression under Edge of Stability Regime
di: Zhang, Peiyuan, et al.
Pubblicazione: (2024)
di: Zhang, Peiyuan, et al.
Pubblicazione: (2024)
Convergence Rates for Gradient Descent on the Edge of Stability in Overparametrised Least Squares
di: MacDonald, Lachlan Ewen, et al.
Pubblicazione: (2025)
di: MacDonald, Lachlan Ewen, et al.
Pubblicazione: (2025)
A Precise Characterization of SGD Stability Using Loss Surface Geometry
di: Dexter, Gregory, et al.
Pubblicazione: (2024)
di: Dexter, Gregory, et al.
Pubblicazione: (2024)
A Rod Flow Model for Adam at the Edge of Stability
di: Regis, Eric, et al.
Pubblicazione: (2026)
di: Regis, Eric, et al.
Pubblicazione: (2026)
Demystifying SGD with Doubly Stochastic Gradients
di: Kim, Kyurae, et al.
Pubblicazione: (2024)
di: Kim, Kyurae, et al.
Pubblicazione: (2024)
SLowcal-SGD: Slow Query Points Improve Local-SGD for Stochastic Convex Optimization
di: Dahan, Tehila, et al.
Pubblicazione: (2023)
di: Dahan, Tehila, et al.
Pubblicazione: (2023)
On the Provable Suboptimality of Momentum SGD in Nonstationary Stochastic Optimization
di: Sahu, Sharan, et al.
Pubblicazione: (2026)
di: Sahu, Sharan, et al.
Pubblicazione: (2026)
A Hessian-Aware Stochastic Differential Equation for Modelling SGD
di: Li, Xiang, et al.
Pubblicazione: (2024)
di: Li, Xiang, et al.
Pubblicazione: (2024)
Rod Flow: A Continuous-Time Model for Gradient Descent at the Edge of Stability
di: Regis, Eric, et al.
Pubblicazione: (2026)
di: Regis, Eric, et al.
Pubblicazione: (2026)
Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling
di: Srećković, Teodora, et al.
Pubblicazione: (2025)
di: Srećković, Teodora, et al.
Pubblicazione: (2025)
SketchySGD: Reliable Stochastic Optimization via Randomized Curvature Estimates
di: Frangella, Zachary, et al.
Pubblicazione: (2022)
di: Frangella, Zachary, et al.
Pubblicazione: (2022)
Enhancing Stochastic Optimization for Statistical Efficiency Using ROOT-SGD with Diminishing Stepsize
di: Li, Chris Junchi
Pubblicazione: (2024)
di: Li, Chris Junchi
Pubblicazione: (2024)
$μ^2$-SGD: Stable Stochastic Optimization via a Double Momentum Mechanism
di: Dahan, Tehila, et al.
Pubblicazione: (2023)
di: Dahan, Tehila, et al.
Pubblicazione: (2023)
Revisiting Gradient Normalization and Clipping for Nonconvex SGD under Heavy-Tailed Noise: Necessity, Sufficiency, and Acceleration
di: Sun, Tao, et al.
Pubblicazione: (2024)
di: Sun, Tao, et al.
Pubblicazione: (2024)
AutoSGD: Automatic Learning Rate Selection for Stochastic Gradient Descent
di: Surjanovic, Nikola, et al.
Pubblicazione: (2025)
di: Surjanovic, Nikola, et al.
Pubblicazione: (2025)
StoSignSGD: Unbiased Structural Stochasticity Fixes SignSGD for Training Large Language Models
di: Yu, Dingzhi, et al.
Pubblicazione: (2026)
di: Yu, Dingzhi, et al.
Pubblicazione: (2026)
Stability and Generalization for Stochastic Recursive Momentum-based Algorithms for (Strongly-)Convex One to $K$-Level Stochastic Optimizations
di: Pan, Xiaokang, et al.
Pubblicazione: (2024)
di: Pan, Xiaokang, et al.
Pubblicazione: (2024)
Provable Complexity Improvement of AdaGrad over SGD: Upper and Lower Bounds in Stochastic Non-Convex Optimization
di: Jiang, Ruichen, et al.
Pubblicazione: (2024)
di: Jiang, Ruichen, et al.
Pubblicazione: (2024)
Making SGD Parameter-Free
di: Carmon, Yair, et al.
Pubblicazione: (2022)
di: Carmon, Yair, et al.
Pubblicazione: (2022)
A Non-Asymptotic Theory of Seminorm Lyapunov Stability: From Deterministic to Stochastic Iterative Algorithms
di: Chen, Zaiwei, et al.
Pubblicazione: (2025)
di: Chen, Zaiwei, et al.
Pubblicazione: (2025)
Shadowheart SGD: Distributed Asynchronous SGD with Optimal Time Complexity Under Arbitrary Computation and Communication Heterogeneity
di: Tyurin, Alexander, et al.
Pubblicazione: (2024)
di: Tyurin, Alexander, et al.
Pubblicazione: (2024)
From PowerSGD to PowerSGD+: Low-Rank Gradient Compression for Distributed Optimization with Convergence Guarantees
di: Xie, Shengping, et al.
Pubblicazione: (2025)
di: Xie, Shengping, et al.
Pubblicazione: (2025)
Controlling the Flow: Stability and Convergence for Stochastic Gradient Descent with Decaying Regularization
di: Kassing, Sebastian, et al.
Pubblicazione: (2025)
di: Kassing, Sebastian, et al.
Pubblicazione: (2025)
Dimension-adapted Momentum Outscales SGD
di: Ferbach, Damien, et al.
Pubblicazione: (2025)
di: Ferbach, Damien, et al.
Pubblicazione: (2025)
Heavy-Tail Phenomenon in Decentralized SGD
di: Gurbuzbalaban, Mert, et al.
Pubblicazione: (2022)
di: Gurbuzbalaban, Mert, et al.
Pubblicazione: (2022)
Revisiting the Last-Iterate Convergence of Stochastic Gradient Methods
di: Liu, Zijian, et al.
Pubblicazione: (2023)
di: Liu, Zijian, et al.
Pubblicazione: (2023)
Algorithmic Stability of Stochastic Gradient Descent with Momentum under Heavy-Tailed Noise
di: Dang, Thanh, et al.
Pubblicazione: (2025)
di: Dang, Thanh, et al.
Pubblicazione: (2025)
SGD with memory: fundamental properties and stochastic acceleration
di: Yarotsky, Dmitry, et al.
Pubblicazione: (2024)
di: Yarotsky, Dmitry, et al.
Pubblicazione: (2024)
Does SGD really happen in tiny subspaces?
di: Song, Minhak, et al.
Pubblicazione: (2024)
di: Song, Minhak, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Momentum Further Constrains Sharpness at the Edge of Stochastic Stability
di: Andreyev, Arseniy, et al.
Pubblicazione: (2026) -
On the Trajectories of SGD Without Replacement
di: Beneventano, Pierfrancesco
Pubblicazione: (2023) -
How Neural Networks Learn the Support is an Implicit Regularization Effect of SGD
di: Beneventano, Pierfrancesco, et al.
Pubblicazione: (2024) -
SGD at the Edge of Stability: The Stochastic Sharpness Gap
di: Liao, Fangshuo, et al.
Pubblicazione: (2026) -
Does SGD Seek Flatness or Sharpness? An Exactly Solvable Model
di: Xu, Yizhou, et al.
Pubblicazione: (2026)