Does SGD Seek Flatness or Sharpness? An Exactly Solvable Model
Fuente:
arXiv
Salvato in:
| Autori principali: | Xu, Yizhou, Beneventano, Pierfrancesco, Chuang, Isaac, Ziyin, Liu |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
On the Trajectories of SGD Without Replacement
di: Beneventano, Pierfrancesco
Pubblicazione: (2023)
di: Beneventano, Pierfrancesco
Pubblicazione: (2023)
Edge of Stochastic Stability: Revisiting the Edge of Stability for SGD
di: Andreyev, Arseniy, et al.
Pubblicazione: (2024)
di: Andreyev, Arseniy, et al.
Pubblicazione: (2024)
How Neural Networks Learn the Support is an Implicit Regularization Effect of SGD
di: Beneventano, Pierfrancesco, et al.
Pubblicazione: (2024)
di: Beneventano, Pierfrancesco, et al.
Pubblicazione: (2024)
Gradient Descent Converges Linearly to Flatter Minima than Gradient Flow in Shallow Linear Networks
di: Beneventano, Pierfrancesco, et al.
Pubblicazione: (2025)
di: Beneventano, Pierfrancesco, et al.
Pubblicazione: (2025)
Does Weight Decay Enhance Training Stability?
di: Saether, Marius, et al.
Pubblicazione: (2026)
di: Saether, Marius, et al.
Pubblicazione: (2026)
Too Sharp, Too Sure: When Calibration Follows Curvature
di: Morosini, Alessandro, et al.
Pubblicazione: (2026)
di: Morosini, Alessandro, et al.
Pubblicazione: (2026)
Momentum Further Constrains Sharpness at the Edge of Stochastic Stability
di: Andreyev, Arseniy, et al.
Pubblicazione: (2026)
di: Andreyev, Arseniy, et al.
Pubblicazione: (2026)
Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias
di: Das, Mohua, et al.
Pubblicazione: (2026)
di: Das, Mohua, et al.
Pubblicazione: (2026)
Does SGD really happen in tiny subspaces?
di: Song, Minhak, et al.
Pubblicazione: (2024)
di: Song, Minhak, et al.
Pubblicazione: (2024)
ROOT-SGD: Sharp Nonasymptotics and Near-Optimal Asymptotics in a Single Algorithm
di: Li, Chris Junchi, et al.
Pubblicazione: (2020)
di: Li, Chris Junchi, et al.
Pubblicazione: (2020)
SGD at the Edge of Stability: The Stochastic Sharpness Gap
di: Liao, Fangshuo, et al.
Pubblicazione: (2026)
di: Liao, Fangshuo, et al.
Pubblicazione: (2026)
Sharp High-Probability Rates for Nonlinear SGD under Heavy-Tailed Noise via Symmetrization
di: Armacki, Aleksandar, et al.
Pubblicazione: (2025)
di: Armacki, Aleksandar, et al.
Pubblicazione: (2025)
Scaling Laws of SignSGD in Linear Regression: When Does It Outperform SGD?
di: Kim, Jihwan, et al.
Pubblicazione: (2026)
di: Kim, Jihwan, et al.
Pubblicazione: (2026)
Schrödinger Bridge with Quadratic State Cost is Exactly Solvable
di: Teter, Alexis M. H., et al.
Pubblicazione: (2024)
di: Teter, Alexis M. H., et al.
Pubblicazione: (2024)
Does Worst-Performing Agent Lead the Pack? Analyzing Agent Dynamics in Unified Distributed SGD
di: Hu, Jie, et al.
Pubblicazione: (2024)
di: Hu, Jie, et al.
Pubblicazione: (2024)
Weyl Calculus and Exactly Solvable Schrödinger Bridges with Quadratic State Cost
di: Teter, Alexis M. H., et al.
Pubblicazione: (2024)
di: Teter, Alexis M. H., et al.
Pubblicazione: (2024)
Faster Convergence of Local SGD for Over-Parameterized Models
di: Qin, Tiancheng, et al.
Pubblicazione: (2022)
di: Qin, Tiancheng, et al.
Pubblicazione: (2022)
SLowcal-SGD: Slow Query Points Improve Local-SGD for Stochastic Convex Optimization
di: Dahan, Tehila, et al.
Pubblicazione: (2023)
di: Dahan, Tehila, et al.
Pubblicazione: (2023)
Making SGD Parameter-Free
di: Carmon, Yair, et al.
Pubblicazione: (2022)
di: Carmon, Yair, et al.
Pubblicazione: (2022)
The Optimality of (Accelerated) SGD for High-Dimensional Quadratic Optimization
di: Zhang, Haihan, et al.
Pubblicazione: (2024)
di: Zhang, Haihan, et al.
Pubblicazione: (2024)
A Hessian-Aware Stochastic Differential Equation for Modelling SGD
di: Li, Xiang, et al.
Pubblicazione: (2024)
di: Li, Xiang, et al.
Pubblicazione: (2024)
From PowerSGD to PowerSGD+: Low-Rank Gradient Compression for Distributed Optimization with Convergence Guarantees
di: Xie, Shengping, et al.
Pubblicazione: (2025)
di: Xie, Shengping, et al.
Pubblicazione: (2025)
Shadowheart SGD: Distributed Asynchronous SGD with Optimal Time Complexity Under Arbitrary Computation and Communication Heterogeneity
di: Tyurin, Alexander, et al.
Pubblicazione: (2024)
di: Tyurin, Alexander, et al.
Pubblicazione: (2024)
StoSignSGD: Unbiased Structural Stochasticity Fixes SignSGD for Training Large Language Models
di: Yu, Dingzhi, et al.
Pubblicazione: (2026)
di: Yu, Dingzhi, et al.
Pubblicazione: (2026)
Dimension-adapted Momentum Outscales SGD
di: Ferbach, Damien, et al.
Pubblicazione: (2025)
di: Ferbach, Damien, et al.
Pubblicazione: (2025)
Heavy-Tail Phenomenon in Decentralized SGD
di: Gurbuzbalaban, Mert, et al.
Pubblicazione: (2022)
di: Gurbuzbalaban, Mert, et al.
Pubblicazione: (2022)
Demystifying SGD with Doubly Stochastic Gradients
di: Kim, Kyurae, et al.
Pubblicazione: (2024)
di: Kim, Kyurae, et al.
Pubblicazione: (2024)
Parameter Symmetry and Noise Equilibrium of Stochastic Gradient Descent
di: Ziyin, Liu, et al.
Pubblicazione: (2024)
di: Ziyin, Liu, et al.
Pubblicazione: (2024)
The Rich and the Simple: On the Implicit Bias of Adam and SGD
di: Vasudeva, Bhavya, et al.
Pubblicazione: (2025)
di: Vasudeva, Bhavya, et al.
Pubblicazione: (2025)
Sign-SGD via Parameter-Free Optimization
di: Medyakov, Daniil, et al.
Pubblicazione: (2025)
di: Medyakov, Daniil, et al.
Pubblicazione: (2025)
Can SGD Handle Heavy-Tailed Noise?
di: Fatkhullin, Ilyas, et al.
Pubblicazione: (2025)
di: Fatkhullin, Ilyas, et al.
Pubblicazione: (2025)
SGD with memory: fundamental properties and stochastic acceleration
di: Yarotsky, Dmitry, et al.
Pubblicazione: (2024)
di: Yarotsky, Dmitry, et al.
Pubblicazione: (2024)
Diagonalisation SGD: Fast & Convergent SGD for Non-Differentiable Models via Reparameterisation and Smoothing
di: Wagner, Dominik, et al.
Pubblicazione: (2024)
di: Wagner, Dominik, et al.
Pubblicazione: (2024)
Byzantine-Robust Distributed SGD: A Unified Analysis and Tight Error Bounds
di: Ruan, Boyuan, et al.
Pubblicazione: (2026)
di: Ruan, Boyuan, et al.
Pubblicazione: (2026)
Catapults in SGD: spikes in the training loss and their impact on generalization through feature learning
di: Zhu, Libin, et al.
Pubblicazione: (2023)
di: Zhu, Libin, et al.
Pubblicazione: (2023)
Phases of Muon: When Muon Eclipses SignSGD
di: Paquette, Elliot, et al.
Pubblicazione: (2026)
di: Paquette, Elliot, et al.
Pubblicazione: (2026)
Optimal Projection-Free Adaptive SGD for Matrix Optimization
di: Kovalev, Dmitry
Pubblicazione: (2026)
di: Kovalev, Dmitry
Pubblicazione: (2026)
On the Provable Suboptimality of Momentum SGD in Nonstationary Stochastic Optimization
di: Sahu, Sharan, et al.
Pubblicazione: (2026)
di: Sahu, Sharan, et al.
Pubblicazione: (2026)
Accelerating Single-Pass SGD for Generalized Linear Prediction
di: Chen, Qian, et al.
Pubblicazione: (2026)
di: Chen, Qian, et al.
Pubblicazione: (2026)
From Gradient Clipping to Normalization for Heavy Tailed SGD
di: Hübler, Florian, et al.
Pubblicazione: (2024)
di: Hübler, Florian, et al.
Pubblicazione: (2024)
Documenti analoghi
-
On the Trajectories of SGD Without Replacement
di: Beneventano, Pierfrancesco
Pubblicazione: (2023) -
Edge of Stochastic Stability: Revisiting the Edge of Stability for SGD
di: Andreyev, Arseniy, et al.
Pubblicazione: (2024) -
How Neural Networks Learn the Support is an Implicit Regularization Effect of SGD
di: Beneventano, Pierfrancesco, et al.
Pubblicazione: (2024) -
Gradient Descent Converges Linearly to Flatter Minima than Gradient Flow in Shallow Linear Networks
di: Beneventano, Pierfrancesco, et al.
Pubblicazione: (2025) -
Does Weight Decay Enhance Training Stability?
di: Saether, Marius, et al.
Pubblicazione: (2026)