Momentum Further Constrains Sharpness at the Edge of Stochastic Stability
Fuente:
arXiv
Saved in:
| Main Authors: | Andreyev, Arseniy, Ananthkumar, Advikar, Walden, Marc, Poggio, Tomaso, Beneventano, Pierfrancesco |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Edge of Stochastic Stability: Revisiting the Edge of Stability for SGD
by: Andreyev, Arseniy, et al.
Published: (2024)
by: Andreyev, Arseniy, et al.
Published: (2024)
Too Sharp, Too Sure: When Calibration Follows Curvature
by: Morosini, Alessandro, et al.
Published: (2026)
by: Morosini, Alessandro, et al.
Published: (2026)
How Neural Networks Learn the Support is an Implicit Regularization Effect of SGD
by: Beneventano, Pierfrancesco, et al.
Published: (2024)
by: Beneventano, Pierfrancesco, et al.
Published: (2024)
Does Weight Decay Enhance Training Stability?
by: Saether, Marius, et al.
Published: (2026)
by: Saether, Marius, et al.
Published: (2026)
Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias
by: Das, Mohua, et al.
Published: (2026)
by: Das, Mohua, et al.
Published: (2026)
On the Trajectories of SGD Without Replacement
by: Beneventano, Pierfrancesco
Published: (2023)
by: Beneventano, Pierfrancesco
Published: (2023)
Zeroth-Order Optimization at the Edge of Stability
by: Song, Minhak, et al.
Published: (2026)
by: Song, Minhak, et al.
Published: (2026)
Does SGD Seek Flatness or Sharpness? An Exactly Solvable Model
by: Xu, Yizhou, et al.
Published: (2026)
by: Xu, Yizhou, et al.
Published: (2026)
Gradient Descent Converges Linearly to Flatter Minima than Gradient Flow in Shallow Linear Networks
by: Beneventano, Pierfrancesco, et al.
Published: (2025)
by: Beneventano, Pierfrancesco, et al.
Published: (2025)
pAI/MSc: ML Theory Research with Humans on the Loop
by: Abdelmoneum, Mahmoud, et al.
Published: (2026)
by: Abdelmoneum, Mahmoud, et al.
Published: (2026)
Stability properties of gradient flow dynamics for the symmetric low-rank matrix factorization problem
by: Mohammadi, Hesameddin, et al.
Published: (2024)
by: Mohammadi, Hesameddin, et al.
Published: (2024)
Stability-Certified Learning of Control Systems with Quadratic Nonlinearities
by: Duff, Igor Pontes, et al.
Published: (2024)
by: Duff, Igor Pontes, et al.
Published: (2024)
Regularity and Stability Properties of Selective SSMs with Discontinuous Gating
by: Zubić, Nikola, et al.
Published: (2025)
by: Zubić, Nikola, et al.
Published: (2025)
Safe and Robust Domains of Attraction for Discrete-Time Systems: A Set-Based Characterization and Certifiable Neural Network Estimation
by: Serry, Mohamed, et al.
Published: (2026)
by: Serry, Mohamed, et al.
Published: (2026)
SINDy-RL: Interpretable and Efficient Model-Based Reinforcement Learning
by: Zolman, Nicholas, et al.
Published: (2024)
by: Zolman, Nicholas, et al.
Published: (2024)
Data-driven Nonlinear Model Reduction using Koopman Theory: Integrated Control Form and NMPC Case Study
by: Schulze, Jan C., et al.
Published: (2024)
by: Schulze, Jan C., et al.
Published: (2024)
Near-Optimal Distributed Linear-Quadratic Regulator for Networked Systems
by: Shin, Sungho, et al.
Published: (2022)
by: Shin, Sungho, et al.
Published: (2022)
Learning Dissipative Neural Dynamical Systems
by: Xu, Yuezhu, et al.
Published: (2023)
by: Xu, Yuezhu, et al.
Published: (2023)
Tradeoffs between convergence rate and noise amplification for momentum-based accelerated optimization algorithms
by: Mohammadi, Hesameddin, et al.
Published: (2022)
by: Mohammadi, Hesameddin, et al.
Published: (2022)
Observability conditions for neural state-space models with eigenvalues and their roots of unity
by: Gracyk, Andrew
Published: (2025)
by: Gracyk, Andrew
Published: (2025)
Safely Learning Dynamical Systems
by: Ahmadi, Amir Ali, et al.
Published: (2023)
by: Ahmadi, Amir Ali, et al.
Published: (2023)
From exponential to finite/fixed-time stability: Applications to optimization
by: Ozaslan, Ibrahim K., et al.
Published: (2024)
by: Ozaslan, Ibrahim K., et al.
Published: (2024)
Neural Operators for Predictor Feedback Control of Nonlinear Delay Systems
by: Bhan, Luke, et al.
Published: (2024)
by: Bhan, Luke, et al.
Published: (2024)
Identifiability of Differential-Algebraic Systems
by: Montanari, Arthur N., et al.
Published: (2024)
by: Montanari, Arthur N., et al.
Published: (2024)
SGD at the Edge of Stability: The Stochastic Sharpness Gap
by: Liao, Fangshuo, et al.
Published: (2026)
by: Liao, Fangshuo, et al.
Published: (2026)
The Global Empirical NTK: Self-Referential Bias and Dimensionality of Gradient Descent Learning
by: Hazelden, James, et al.
Published: (2026)
by: Hazelden, James, et al.
Published: (2026)
Convergence of optimizers implies eigenvalues filtering at equilibrium
by: Bolte, Jerome, et al.
Published: (2025)
by: Bolte, Jerome, et al.
Published: (2025)
Circular Microalgae-Based Carbon Control for Net Zero
by: Zocco, Federico, et al.
Published: (2025)
by: Zocco, Federico, et al.
Published: (2025)
Near-optimal Closed-loop Method via Lyapunov Damping for Convex Optimization
by: Maier, Severin, et al.
Published: (2023)
by: Maier, Severin, et al.
Published: (2023)
Nonlinear Dynamics In Optimization Landscape of Shallow Neural Networks with Tunable Leaky ReLU
by: Liu, Jingzhou
Published: (2025)
by: Liu, Jingzhou
Published: (2025)
Synchronization on circles and spheres with nonlinear interactions
by: Criscitiello, Christopher, et al.
Published: (2024)
by: Criscitiello, Christopher, et al.
Published: (2024)
On the Convergence of Overlapping Schwarz Decomposition for Nonlinear Optimal Control
by: Na, Sen, et al.
Published: (2020)
by: Na, Sen, et al.
Published: (2020)
Learning Spatio-Temporal Dynamics via Operator-Valued RKHS and Kernel Koopman Methods
by: Withanachchi, Mahishanka
Published: (2025)
by: Withanachchi, Mahishanka
Published: (2025)
Memorization and Regularization in Generative Diffusion Models
by: Baptista, Ricardo, et al.
Published: (2025)
by: Baptista, Ricardo, et al.
Published: (2025)
Kernel Sum of Squares for Data Adapted Kernel Learning of Dynamical Systems from Data: A global optimization approach
by: Lengyel, Daniel, et al.
Published: (2024)
by: Lengyel, Daniel, et al.
Published: (2024)
Iterative regularization in classification via hinge loss diagonal descent
by: Apidopoulos, Vassilis, et al.
Published: (2022)
by: Apidopoulos, Vassilis, et al.
Published: (2022)
Online Optimization and Ambiguity-based Learning of Distributionally Uncertain Dynamic Systems
by: Li, Dan, et al.
Published: (2021)
by: Li, Dan, et al.
Published: (2021)
Safe and Near-Optimal Control with Online Dynamics Learning
by: Prajapat, Manish, et al.
Published: (2025)
by: Prajapat, Manish, et al.
Published: (2025)
Reservoir Predictive Path Integral Control for Unknown Nonlinear Dynamics
by: Inoue, Daisuke, et al.
Published: (2025)
by: Inoue, Daisuke, et al.
Published: (2025)
On Convex Data-Driven Inverse Optimal Control for Nonlinear, Non-stationary and Stochastic Systems
by: Garrabe, Emiland, et al.
Published: (2023)
by: Garrabe, Emiland, et al.
Published: (2023)
Similar Items
-
Edge of Stochastic Stability: Revisiting the Edge of Stability for SGD
by: Andreyev, Arseniy, et al.
Published: (2024) -
Too Sharp, Too Sure: When Calibration Follows Curvature
by: Morosini, Alessandro, et al.
Published: (2026) -
How Neural Networks Learn the Support is an Implicit Regularization Effect of SGD
by: Beneventano, Pierfrancesco, et al.
Published: (2024) -
Does Weight Decay Enhance Training Stability?
by: Saether, Marius, et al.
Published: (2026) -
Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias
by: Das, Mohua, et al.
Published: (2026)