Fundamental Convergence Analysis of Sharpness-Aware Minimization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Khanh, Pham Duy, Luong, Hoang-Chau, Mordukhovich, Boris S., Tran, Dat Ba
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912077432487936
author Khanh, Pham Duy
Luong, Hoang-Chau
Mordukhovich, Boris S.
Tran, Dat Ba
author_facet Khanh, Pham Duy
Luong, Hoang-Chau
Mordukhovich, Boris S.
Tran, Dat Ba
contents The paper investigates the fundamental convergence properties of Sharpness-Aware Minimization (SAM), a recently proposed gradient-based optimization method [Foret et al., 2021] that significantly improves the generalization of deep neural networks. The convergence properties, including the stationarity of accumulation points, the convergence of the sequence of gradients to the origin, the sequence of function values to the optimal value, and the sequence of iterates to the optimal solution, are established for the method. The universality of the provided convergence analysis, based on inexact gradient descent frameworks Khanh et al. [2023b], allows its extensions to efficient normalized versions of SAM such as F-SAM [Li et al., 2024], VaSSO [Li and Giannakis, 2023], RSAM [Liu et al., 2022], and to the unnormalized versions of SAM such as USAM [Andriushchenko and Flammarion, 2022]. Numerical experiments are conducted on classification tasks using deep learning models to confirm the practical aspects of our analysis.
format Preprint
id arxiv_https___arxiv_org_abs_2401_08060
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Fundamental Convergence Analysis of Sharpness-Aware Minimization
Khanh, Pham Duy
Luong, Hoang-Chau
Mordukhovich, Boris S.
Tran, Dat Ba
Optimization and Control
The paper investigates the fundamental convergence properties of Sharpness-Aware Minimization (SAM), a recently proposed gradient-based optimization method [Foret et al., 2021] that significantly improves the generalization of deep neural networks. The convergence properties, including the stationarity of accumulation points, the convergence of the sequence of gradients to the origin, the sequence of function values to the optimal value, and the sequence of iterates to the optimal solution, are established for the method. The universality of the provided convergence analysis, based on inexact gradient descent frameworks Khanh et al. [2023b], allows its extensions to efficient normalized versions of SAM such as F-SAM [Li et al., 2024], VaSSO [Li and Giannakis, 2023], RSAM [Liu et al., 2022], and to the unnormalized versions of SAM such as USAM [Andriushchenko and Flammarion, 2022]. Numerical experiments are conducted on classification tasks using deep learning models to confirm the practical aspects of our analysis.
title Fundamental Convergence Analysis of Sharpness-Aware Minimization
topic Optimization and Control
url https://arxiv.org/abs/2401.08060