Sharpness-Aware Minimization Enhances Feature Quality via Balanced Learning

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Springer, Jacob Mitchell, Nagarajan, Vaishnavh, Raghunathan, Aditi
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914816968359936
author Springer, Jacob Mitchell
Nagarajan, Vaishnavh
Raghunathan, Aditi
author_facet Springer, Jacob Mitchell
Nagarajan, Vaishnavh
Raghunathan, Aditi
contents Sharpness-Aware Minimization (SAM) has emerged as a promising alternative optimizer to stochastic gradient descent (SGD). The originally-proposed motivation behind SAM was to bias neural networks towards flatter minima that are believed to generalize better. However, recent studies have shown conflicting evidence on the relationship between flatness and generalization, suggesting that flatness does fully explain SAM's success. Sidestepping this debate, we identify an orthogonal effect of SAM that is beneficial out-of-distribution: we argue that SAM implicitly balances the quality of diverse features. SAM achieves this effect by adaptively suppressing well-learned features which gives remaining features opportunity to be learned. We show that this mechanism is beneficial in datasets that contain redundant or spurious features where SGD falls for the simplicity bias and would not otherwise learn all available features. Our insights are supported by experiments on real data: we demonstrate that SAM improves the quality of features in datasets containing redundant or spurious features, including CelebA, Waterbirds, CIFAR-MNIST, and DomainBed.
format Preprint
id arxiv_https___arxiv_org_abs_2405_20439
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Sharpness-Aware Minimization Enhances Feature Quality via Balanced Learning
Springer, Jacob Mitchell
Nagarajan, Vaishnavh
Raghunathan, Aditi
Machine Learning
Sharpness-Aware Minimization (SAM) has emerged as a promising alternative optimizer to stochastic gradient descent (SGD). The originally-proposed motivation behind SAM was to bias neural networks towards flatter minima that are believed to generalize better. However, recent studies have shown conflicting evidence on the relationship between flatness and generalization, suggesting that flatness does fully explain SAM's success. Sidestepping this debate, we identify an orthogonal effect of SAM that is beneficial out-of-distribution: we argue that SAM implicitly balances the quality of diverse features. SAM achieves this effect by adaptively suppressing well-learned features which gives remaining features opportunity to be learned. We show that this mechanism is beneficial in datasets that contain redundant or spurious features where SGD falls for the simplicity bias and would not otherwise learn all available features. Our insights are supported by experiments on real data: we demonstrate that SAM improves the quality of features in datasets containing redundant or spurious features, including CelebA, Waterbirds, CIFAR-MNIST, and DomainBed.
title Sharpness-Aware Minimization Enhances Feature Quality via Balanced Learning
topic Machine Learning
url https://arxiv.org/abs/2405.20439