Effective Gradient Sample Size via Variation Estimation for Accelerating Sharpness aware Minimization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Deng, Jiaxin, Pang, Junbiao, Zhang, Baochang, Wang, Tian
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910366759387136
author Deng, Jiaxin
Pang, Junbiao
Zhang, Baochang
Wang, Tian
author_facet Deng, Jiaxin
Pang, Junbiao
Zhang, Baochang
Wang, Tian
contents Sharpness-aware Minimization (SAM) has been proposed recently to improve model generalization ability. However, SAM calculates the gradient twice in each optimization step, thereby doubling the computation costs compared to stochastic gradient descent (SGD). In this paper, we propose a simple yet efficient sampling method to significantly accelerate SAM. Concretely, we discover that the gradient of SAM is a combination of the gradient of SGD and the Projection of the Second-order gradient matrix onto the First-order gradient (PSF). PSF exhibits a gradually increasing frequency of change during the training process. To leverage this observation, we propose an adaptive sampling method based on the variation of PSF, and we reuse the sampled PSF for non-sampling iterations. Extensive empirical results illustrate that the proposed method achieved state-of-the-art accuracies comparable to SAM on diverse network architectures.
format Preprint
id arxiv_https___arxiv_org_abs_2403_08821
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Effective Gradient Sample Size via Variation Estimation for Accelerating Sharpness aware Minimization
Deng, Jiaxin
Pang, Junbiao
Zhang, Baochang
Wang, Tian
Computer Vision and Pattern Recognition
Machine Learning
Sharpness-aware Minimization (SAM) has been proposed recently to improve model generalization ability. However, SAM calculates the gradient twice in each optimization step, thereby doubling the computation costs compared to stochastic gradient descent (SGD). In this paper, we propose a simple yet efficient sampling method to significantly accelerate SAM. Concretely, we discover that the gradient of SAM is a combination of the gradient of SGD and the Projection of the Second-order gradient matrix onto the First-order gradient (PSF). PSF exhibits a gradually increasing frequency of change during the training process. To leverage this observation, we propose an adaptive sampling method based on the variation of PSF, and we reuse the sampled PSF for non-sampling iterations. Extensive empirical results illustrate that the proposed method achieved state-of-the-art accuracies comparable to SAM on diverse network architectures.
title Effective Gradient Sample Size via Variation Estimation for Accelerating Sharpness aware Minimization
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2403.08821