Sharpness-Aware Minimization with Adaptive Regularization for Training Deep Neural Networks

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Zou, Jinping, Deng, Xiaoge, Sun, Tao
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866910759154352128
author Zou, Jinping
Deng, Xiaoge
Sun, Tao
author_facet Zou, Jinping
Deng, Xiaoge
Sun, Tao
contents Sharpness-Aware Minimization (SAM) has proven highly effective in improving model generalization in machine learning tasks. However, SAM employs a fixed hyperparameter associated with the regularization to characterize the sharpness of the model. Despite its success, research on adaptive regularization methods based on SAM remains scarce. In this paper, we propose the SAM with Adaptive Regularization (SAMAR), which introduces a flexible sharpness ratio rule to update the regularization parameter dynamically. We provide theoretical proof of the convergence of SAMAR for functions satisfying the Lipschitz continuity. Additionally, experiments on image recognition tasks using CIFAR-10 and CIFAR-100 demonstrate that SAMAR enhances accuracy and model generalization.
format Preprint
id arxiv_https___arxiv_org_abs_2412_16854
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Sharpness-Aware Minimization with Adaptive Regularization for Training Deep Neural Networks
Zou, Jinping
Deng, Xiaoge
Sun, Tao
Machine Learning
Computer Vision and Pattern Recognition
Sharpness-Aware Minimization (SAM) has proven highly effective in improving model generalization in machine learning tasks. However, SAM employs a fixed hyperparameter associated with the regularization to characterize the sharpness of the model. Despite its success, research on adaptive regularization methods based on SAM remains scarce. In this paper, we propose the SAM with Adaptive Regularization (SAMAR), which introduces a flexible sharpness ratio rule to update the regularization parameter dynamically. We provide theoretical proof of the convergence of SAMAR for functions satisfying the Lipschitz continuity. Additionally, experiments on image recognition tasks using CIFAR-10 and CIFAR-100 demonstrate that SAMAR enhances accuracy and model generalization.
title Sharpness-Aware Minimization with Adaptive Regularization for Training Deep Neural Networks
topic Machine Learning
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.16854