Sharpness-Aware Minimization: General Analysis and Improved Rates
Fuente:
arXiv
Saved in:
| Main Authors: | Oikonomou, Dimitris, Loizou, Nicolas |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Adaptive Sharpness-Aware Minimization with a Polyak-type Step size: A Theory-Grounded Scheduler
by: Oikonomou, Dimitris, et al.
Published: (2026)
by: Oikonomou, Dimitris, et al.
Published: (2026)
Safeguarded Stochastic Polyak Step Sizes for Non-smooth Optimization: Robust Performance Without Small (Sub)Gradients
by: Oikonomou, Dimitris, et al.
Published: (2025)
by: Oikonomou, Dimitris, et al.
Published: (2025)
Stochastic Polyak Step-sizes and Momentum: Convergence Guarantees and Practical Performance
by: Oikonomou, Dimitris, et al.
Published: (2024)
by: Oikonomou, Dimitris, et al.
Published: (2024)
Sharpness-Aware Minimization Can Hallucinate Minimizers
by: Park, Chanwoong, et al.
Published: (2025)
by: Park, Chanwoong, et al.
Published: (2025)
Extragradient Method for $(L_0, L_1)$-Lipschitz Root-finding Problems
by: Choudhury, Sayantan, et al.
Published: (2025)
by: Choudhury, Sayantan, et al.
Published: (2025)
Convergence of Sharpness-Aware Minimization Algorithms using Increasing Batch Size and Decaying Learning Rate
by: Harada, Hinata, et al.
Published: (2024)
by: Harada, Hinata, et al.
Published: (2024)
DGSAM: Domain Generalization via Individual Sharpness-Aware Minimization
by: Song, Youngjun, et al.
Published: (2025)
by: Song, Youngjun, et al.
Published: (2025)
Communication-Efficient Gradient Descent-Accent Methods for Distributed Variational Inequalities: Unified Analysis and Local Updates
by: Zhang, Siqi, et al.
Published: (2023)
by: Zhang, Siqi, et al.
Published: (2023)
Multiplayer Federated Learning: Reaching Equilibrium with Less Communication
by: Yoon, TaeHo, et al.
Published: (2025)
by: Yoon, TaeHo, et al.
Published: (2025)
Locally Adaptive Federated Learning
by: Mukherjee, Sohom, et al.
Published: (2023)
by: Mukherjee, Sohom, et al.
Published: (2023)
Stochastic Extragradient with Random Reshuffling: Improved Convergence for Variational Inequalities
by: Emmanouilidis, Konstantinos, et al.
Published: (2024)
by: Emmanouilidis, Konstantinos, et al.
Published: (2024)
Dissipative Gradient Descent Ascent Method: A Control Theory Inspired Algorithm for Min-max Optimization
by: Zheng, Tianqi, et al.
Published: (2024)
by: Zheng, Tianqi, et al.
Published: (2024)
Critical Influence of Overparameterization on Sharpness-aware Minimization
by: Shin, Sungbin, et al.
Published: (2023)
by: Shin, Sungbin, et al.
Published: (2023)
Adaptive Algorithms with Sharp Convergence Rates for Stochastic Hierarchical Optimization
by: Gong, Xiaochuan, et al.
Published: (2025)
by: Gong, Xiaochuan, et al.
Published: (2025)
On the Duality Between Sharpness-Aware Minimization and Adversarial Training
by: Zhang, Yihao, et al.
Published: (2024)
by: Zhang, Yihao, et al.
Published: (2024)
Non-Euclidean SGD for Structured Optimization: Unified Analysis and Improved Rates
by: Kovalev, Dmitry, et al.
Published: (2025)
by: Kovalev, Dmitry, et al.
Published: (2025)
Sharp High-Probability Rates for Nonlinear SGD under Heavy-Tailed Noise via Symmetrization
by: Armacki, Aleksandar, et al.
Published: (2025)
by: Armacki, Aleksandar, et al.
Published: (2025)
Improved Learning Rates for Stochastic Optimization
by: Li, Shaojie, et al.
Published: (2021)
by: Li, Shaojie, et al.
Published: (2021)
Remove that Square Root: A New Efficient Scale-Invariant Version of AdaGrad
by: Choudhury, Sayantan, et al.
Published: (2024)
by: Choudhury, Sayantan, et al.
Published: (2024)
Dynamic Regularized Sharpness Aware Minimization in Federated Learning: Approaching Global Consistency and Smooth Landscape
by: Sun, Yan, et al.
Published: (2023)
by: Sun, Yan, et al.
Published: (2023)
Improved Convergence Rates of Muon Optimizer for Nonconvex Optimization
by: Nagashima, Shuntaro, et al.
Published: (2026)
by: Nagashima, Shuntaro, et al.
Published: (2026)
Convergence Rate Analysis of LION
by: Dong, Yiming, et al.
Published: (2024)
by: Dong, Yiming, et al.
Published: (2024)
Fundamental Convergence Analysis of Sharpness-Aware Minimization
by: Khanh, Pham Duy, et al.
Published: (2024)
by: Khanh, Pham Duy, et al.
Published: (2024)
Learning Rate Annealing Improves Tuning Robustness in Stochastic Optimization
by: Attia, Amit, et al.
Published: (2025)
by: Attia, Amit, et al.
Published: (2025)
Sharpness of Minima in Deep Matrix Factorization
by: Kamber, Anil, et al.
Published: (2025)
by: Kamber, Anil, et al.
Published: (2025)
From Data to Uncertainty Sets: a Machine Learning Approach
by: Bertsimas, Dimitris, et al.
Published: (2025)
by: Bertsimas, Dimitris, et al.
Published: (2025)
Overfitting in Adaptive Robust Optimization
by: Zhu, Karl, et al.
Published: (2025)
by: Zhu, Karl, et al.
Published: (2025)
Global Optimization: A Machine Learning Approach
by: Bertsimas, Dimitris, et al.
Published: (2023)
by: Bertsimas, Dimitris, et al.
Published: (2023)
Catastrophe Insurance: An Adaptive Robust Optimization Approach
by: Bertsimas, Dimitris, et al.
Published: (2024)
by: Bertsimas, Dimitris, et al.
Published: (2024)
Robust Regression over Averaged Uncertainty
by: Bertsimas, Dimitris, et al.
Published: (2023)
by: Bertsimas, Dimitris, et al.
Published: (2023)
Using Taylor-Approximated Gradients to Improve the Frank-Wolfe Method for Empirical Risk Minimization
by: Xiong, Zikai, et al.
Published: (2022)
by: Xiong, Zikai, et al.
Published: (2022)
Empirical Risk Minimization with Shuffled SGD: A Primal-Dual Perspective and Improved Bounds
by: Cai, Xufeng, et al.
Published: (2023)
by: Cai, Xufeng, et al.
Published: (2023)
Adaptive Batch Size and Learning Rate Scheduler for Stochastic Gradient Descent Based on Minimization of Stochastic First-order Oracle Complexity
by: Umeda, Hikaru, et al.
Published: (2025)
by: Umeda, Hikaru, et al.
Published: (2025)
A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies
by: Nanda, Phalguni, et al.
Published: (2025)
by: Nanda, Phalguni, et al.
Published: (2025)
A Machine Learning Approach to Two-Stage Adaptive Robust Optimization
by: Bertsimas, Dimitris, et al.
Published: (2023)
by: Bertsimas, Dimitris, et al.
Published: (2023)
Lotka-Sharpe Neural Operators for Control of Population PDEs
by: Krstic, Miroslav, et al.
Published: (2026)
by: Krstic, Miroslav, et al.
Published: (2026)
Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence
by: Liu, Yuxing, et al.
Published: (2025)
by: Liu, Yuxing, et al.
Published: (2025)
Does SGD Seek Flatness or Sharpness? An Exactly Solvable Model
by: Xu, Yizhou, et al.
Published: (2026)
by: Xu, Yizhou, et al.
Published: (2026)
Sharp Global Guarantees for Nonconvex Low-rank Recovery in the Noisy Overparameterized Regime
by: Zhang, Richard Y.
Published: (2021)
by: Zhang, Richard Y.
Published: (2021)
ROOT-SGD: Sharp Nonasymptotics and Near-Optimal Asymptotics in a Single Algorithm
by: Li, Chris Junchi, et al.
Published: (2020)
by: Li, Chris Junchi, et al.
Published: (2020)
Similar Items
-
Adaptive Sharpness-Aware Minimization with a Polyak-type Step size: A Theory-Grounded Scheduler
by: Oikonomou, Dimitris, et al.
Published: (2026) -
Safeguarded Stochastic Polyak Step Sizes for Non-smooth Optimization: Robust Performance Without Small (Sub)Gradients
by: Oikonomou, Dimitris, et al.
Published: (2025) -
Stochastic Polyak Step-sizes and Momentum: Convergence Guarantees and Practical Performance
by: Oikonomou, Dimitris, et al.
Published: (2024) -
Sharpness-Aware Minimization Can Hallucinate Minimizers
by: Park, Chanwoong, et al.
Published: (2025) -
Extragradient Method for $(L_0, L_1)$-Lipschitz Root-finding Problems
by: Choudhury, Sayantan, et al.
Published: (2025)