Sharpness-Aware Minimization Revisited: Weighted Sharpness as a Regularization Term

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yue, Yun, Jiang, Jiadi, Ye, Zhiling, Gao, Ning, Liu, Yongchao, Zhang, Ke
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912144012869632
author Yue, Yun
Jiang, Jiadi
Ye, Zhiling
Gao, Ning
Liu, Yongchao
Zhang, Ke
author_facet Yue, Yun
Jiang, Jiadi
Ye, Zhiling
Gao, Ning
Liu, Yongchao
Zhang, Ke
contents Deep Neural Networks (DNNs) generalization is known to be closely related to the flatness of minima, leading to the development of Sharpness-Aware Minimization (SAM) for seeking flatter minima and better generalization. In this paper, we revisit the loss of SAM and propose a more general method, called WSAM, by incorporating sharpness as a regularization term. We prove its generalization bound through the combination of PAC and Bayes-PAC techniques, and evaluate its performance on various public datasets. The results demonstrate that WSAM achieves improved generalization, or is at least highly competitive, compared to the vanilla optimizer, SAM and its variants. The code is available at https://github.com/intelligent-machine-learning/atorch/tree/main/atorch/optimizers.
format Preprint
id arxiv_https___arxiv_org_abs_2305_15817
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Sharpness-Aware Minimization Revisited: Weighted Sharpness as a Regularization Term
Yue, Yun
Jiang, Jiadi
Ye, Zhiling
Gao, Ning
Liu, Yongchao
Zhang, Ke
Machine Learning
Deep Neural Networks (DNNs) generalization is known to be closely related to the flatness of minima, leading to the development of Sharpness-Aware Minimization (SAM) for seeking flatter minima and better generalization. In this paper, we revisit the loss of SAM and propose a more general method, called WSAM, by incorporating sharpness as a regularization term. We prove its generalization bound through the combination of PAC and Bayes-PAC techniques, and evaluate its performance on various public datasets. The results demonstrate that WSAM achieves improved generalization, or is at least highly competitive, compared to the vanilla optimizer, SAM and its variants. The code is available at https://github.com/intelligent-machine-learning/atorch/tree/main/atorch/optimizers.
title Sharpness-Aware Minimization Revisited: Weighted Sharpness as a Regularization Term
topic Machine Learning
url https://arxiv.org/abs/2305.15817