Adaptive Sharpness-Aware Minimization with a Polyak-type Step size: A Theory-Grounded Scheduler
Fuente:
arXiv
Saved in:
| Main Authors: | Oikonomou, Dimitris, Loizou, Nicolas |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Stochastic Polyak Step-sizes and Momentum: Convergence Guarantees and Practical Performance
by: Oikonomou, Dimitris, et al.
Published: (2024)
by: Oikonomou, Dimitris, et al.
Published: (2024)
Sharpness-Aware Minimization: General Analysis and Improved Rates
by: Oikonomou, Dimitris, et al.
Published: (2025)
by: Oikonomou, Dimitris, et al.
Published: (2025)
Safeguarded Stochastic Polyak Step Sizes for Non-smooth Optimization: Robust Performance Without Small (Sub)Gradients
by: Oikonomou, Dimitris, et al.
Published: (2025)
by: Oikonomou, Dimitris, et al.
Published: (2025)
Taking the Road Less Scheduled with Adaptive Polyak Steps
by: Oikonomou, Dimitris, et al.
Published: (2025)
by: Oikonomou, Dimitris, et al.
Published: (2025)
Locally Adaptive Federated Learning
by: Mukherjee, Sohom, et al.
Published: (2023)
by: Mukherjee, Sohom, et al.
Published: (2023)
Sharpness-Aware Minimization Can Hallucinate Minimizers
by: Park, Chanwoong, et al.
Published: (2025)
by: Park, Chanwoong, et al.
Published: (2025)
Constrained Online Convex Optimization with Polyak Feasibility Steps
by: Hutchinson, Spencer, et al.
Published: (2025)
by: Hutchinson, Spencer, et al.
Published: (2025)
Extragradient Method for $(L_0, L_1)$-Lipschitz Root-finding Problems
by: Choudhury, Sayantan, et al.
Published: (2025)
by: Choudhury, Sayantan, et al.
Published: (2025)
Dissipative Gradient Descent Ascent Method: A Control Theory Inspired Algorithm for Min-max Optimization
by: Zheng, Tianqi, et al.
Published: (2024)
by: Zheng, Tianqi, et al.
Published: (2024)
Dynamics of SGD with Stochastic Polyak Stepsizes: Truly Adaptive Variants and Convergence to Exact Solution
by: Orvieto, Antonio, et al.
Published: (2022)
by: Orvieto, Antonio, et al.
Published: (2022)
Sparse Polyak: an adaptive step size rule for high-dimensional M-estimation
by: Qiao, Tianqi, et al.
Published: (2025)
by: Qiao, Tianqi, et al.
Published: (2025)
SANIA: Polyak-type Optimization Framework Leads to Scale Invariant Stochastic Algorithms
by: Abdukhakimov, Farshed, et al.
Published: (2023)
by: Abdukhakimov, Farshed, et al.
Published: (2023)
Adaptive SGD with Line-Search and Polyak Stepsizes: Nonconvex Convergence and Accelerated Rates
by: Wu, Haotian
Published: (2025)
by: Wu, Haotian
Published: (2025)
Multiplayer Federated Learning: Reaching Equilibrium with Less Communication
by: Yoon, TaeHo, et al.
Published: (2025)
by: Yoon, TaeHo, et al.
Published: (2025)
DGSAM: Domain Generalization via Individual Sharpness-Aware Minimization
by: Song, Youngjun, et al.
Published: (2025)
by: Song, Youngjun, et al.
Published: (2025)
Convergence of Sharpness-Aware Minimization Algorithms using Increasing Batch Size and Decaying Learning Rate
by: Harada, Hinata, et al.
Published: (2024)
by: Harada, Hinata, et al.
Published: (2024)
Critical Influence of Overparameterization on Sharpness-aware Minimization
by: Shin, Sungbin, et al.
Published: (2023)
by: Shin, Sungbin, et al.
Published: (2023)
Communication-Efficient Gradient Descent-Accent Methods for Distributed Variational Inequalities: Unified Analysis and Local Updates
by: Zhang, Siqi, et al.
Published: (2023)
by: Zhang, Siqi, et al.
Published: (2023)
Non-convex Stochastic Composite Optimization with Polyak Momentum
by: Gao, Yuan, et al.
Published: (2024)
by: Gao, Yuan, et al.
Published: (2024)
Parameter-free Clipped Gradient Descent Meets Polyak
by: Takezawa, Yuki, et al.
Published: (2024)
by: Takezawa, Yuki, et al.
Published: (2024)
New Perspectives on the Polyak Stepsize: Surrogate Functions and Negative Results
by: Orabona, Francesco, et al.
Published: (2025)
by: Orabona, Francesco, et al.
Published: (2025)
Overfitting in Adaptive Robust Optimization
by: Zhu, Karl, et al.
Published: (2025)
by: Zhu, Karl, et al.
Published: (2025)
Faster Stochastic Algorithms for Minimax Optimization under Polyak--Łojasiewicz Conditions
by: Chen, Lesi, et al.
Published: (2023)
by: Chen, Lesi, et al.
Published: (2023)
Sparse Polyak with optimal thresholding operators for high-dimensional M-estimation
by: Qiao, Tianqi, et al.
Published: (2025)
by: Qiao, Tianqi, et al.
Published: (2025)
Minimisation of Polyak-Łojasewicz Functions Using Random Zeroth-Order Oracles
by: Farzin, Amir Ali, et al.
Published: (2024)
by: Farzin, Amir Ali, et al.
Published: (2024)
On the Complexity of Finite-Sum Smooth Optimization under the Polyak-Łojasiewicz Condition
by: Bai, Yunyan, et al.
Published: (2024)
by: Bai, Yunyan, et al.
Published: (2024)
Catastrophe Insurance: An Adaptive Robust Optimization Approach
by: Bertsimas, Dimitris, et al.
Published: (2024)
by: Bertsimas, Dimitris, et al.
Published: (2024)
A Local Polyak-Lojasiewicz and Descent Lemma of Gradient Descent For Overparametrized Linear Models
by: Xu, Ziqing, et al.
Published: (2025)
by: Xu, Ziqing, et al.
Published: (2025)
A Machine Learning Approach to Two-Stage Adaptive Robust Optimization
by: Bertsimas, Dimitris, et al.
Published: (2023)
by: Bertsimas, Dimitris, et al.
Published: (2023)
On the Duality Between Sharpness-Aware Minimization and Adversarial Training
by: Zhang, Yihao, et al.
Published: (2024)
by: Zhang, Yihao, et al.
Published: (2024)
Gradient Descent with Polyak's Momentum Finds Flatter Minima via Large Catapults
by: Phunyaphibarn, Prin, et al.
Published: (2023)
by: Phunyaphibarn, Prin, et al.
Published: (2023)
High-Probability Bounds for SGD under the Polyak-Lojasiewicz Condition with Markovian Noise
by: Kar, Avik, et al.
Published: (2026)
by: Kar, Avik, et al.
Published: (2026)
Adaptive Batch Size and Learning Rate Scheduler for Stochastic Gradient Descent Based on Minimization of Stochastic First-order Oracle Complexity
by: Umeda, Hikaru, et al.
Published: (2025)
by: Umeda, Hikaru, et al.
Published: (2025)
Adaptive Algorithms with Sharp Convergence Rates for Stochastic Hierarchical Optimization
by: Gong, Xiaochuan, et al.
Published: (2025)
by: Gong, Xiaochuan, et al.
Published: (2025)
Stochastic Extragradient with Random Reshuffling: Improved Convergence for Variational Inequalities
by: Emmanouilidis, Konstantinos, et al.
Published: (2024)
by: Emmanouilidis, Konstantinos, et al.
Published: (2024)
Minimizing the Weighted Number of Tardy Jobs: Data-Driven Heuristic for Single-Machine Scheduling
by: Antonov, Nikolai, et al.
Published: (2025)
by: Antonov, Nikolai, et al.
Published: (2025)
AdaBatchGrad: Combining Adaptive Batch Size and Adaptive Step Size
by: Ostroukhov, Petr, et al.
Published: (2024)
by: Ostroukhov, Petr, et al.
Published: (2024)
Shuffling the Data, Stretching the Step-size: Sharper Bias in constant step-size SGD
by: Emmanouilidis, Konstantinos, et al.
Published: (2026)
by: Emmanouilidis, Konstantinos, et al.
Published: (2026)
Adaptive Step Sizes for Preconditioned Stochastic Gradient Descent
by: Köhne, Frederik, et al.
Published: (2023)
by: Köhne, Frederik, et al.
Published: (2023)
Remove that Square Root: A New Efficient Scale-Invariant Version of AdaGrad
by: Choudhury, Sayantan, et al.
Published: (2024)
by: Choudhury, Sayantan, et al.
Published: (2024)
Similar Items
-
Stochastic Polyak Step-sizes and Momentum: Convergence Guarantees and Practical Performance
by: Oikonomou, Dimitris, et al.
Published: (2024) -
Sharpness-Aware Minimization: General Analysis and Improved Rates
by: Oikonomou, Dimitris, et al.
Published: (2025) -
Safeguarded Stochastic Polyak Step Sizes for Non-smooth Optimization: Robust Performance Without Small (Sub)Gradients
by: Oikonomou, Dimitris, et al.
Published: (2025) -
Taking the Road Less Scheduled with Adaptive Polyak Steps
by: Oikonomou, Dimitris, et al.
Published: (2025) -
Locally Adaptive Federated Learning
by: Mukherjee, Sohom, et al.
Published: (2023)