Improving Generalization and Convergence by Enhancing Implicit Regularization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Mingze, Wang, Jinbo, He, Haotian, Wang, Zilin, Huang, Guanhua, Xiong, Feiyu, Li, Zhiyu, E, Weinan, Wu, Lei
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910679679631360
author Wang, Mingze
Wang, Jinbo
He, Haotian
Wang, Zilin
Huang, Guanhua
Xiong, Feiyu
Li, Zhiyu
E, Weinan
Wu, Lei
author_facet Wang, Mingze
Wang, Jinbo
He, Haotian
Wang, Zilin
Huang, Guanhua
Xiong, Feiyu
Li, Zhiyu
E, Weinan
Wu, Lei
contents In this work, we propose an Implicit Regularization Enhancement (IRE) framework to accelerate the discovery of flat solutions in deep learning, thereby improving generalization and convergence. Specifically, IRE decouples the dynamics of flat and sharp directions, which boosts the sharpness reduction along flat directions while maintaining the training stability in sharp directions. We show that IRE can be practically incorporated with {\em generic base optimizers} without introducing significant computational overload. Experiments show that IRE consistently improves the generalization performance for image classification tasks across a variety of benchmark datasets (CIFAR-10/100, ImageNet) and models (ResNets and ViTs). Surprisingly, IRE also achieves a $2\times$ {\em speed-up} compared to AdamW in the pre-training of Llama models (of sizes ranging from 60M to 229M) on datasets including Wikitext-103, Minipile, and Openwebtext. Moreover, we provide theoretical guarantees, showing that IRE can substantially accelerate the convergence towards flat minima in Sharpness-aware Minimization (SAM).
format Preprint
id arxiv_https___arxiv_org_abs_2405_20763
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Improving Generalization and Convergence by Enhancing Implicit Regularization
Wang, Mingze
Wang, Jinbo
He, Haotian
Wang, Zilin
Huang, Guanhua
Xiong, Feiyu
Li, Zhiyu
E, Weinan
Wu, Lei
Machine Learning
Optimization and Control
In this work, we propose an Implicit Regularization Enhancement (IRE) framework to accelerate the discovery of flat solutions in deep learning, thereby improving generalization and convergence. Specifically, IRE decouples the dynamics of flat and sharp directions, which boosts the sharpness reduction along flat directions while maintaining the training stability in sharp directions. We show that IRE can be practically incorporated with {\em generic base optimizers} without introducing significant computational overload. Experiments show that IRE consistently improves the generalization performance for image classification tasks across a variety of benchmark datasets (CIFAR-10/100, ImageNet) and models (ResNets and ViTs). Surprisingly, IRE also achieves a $2\times$ {\em speed-up} compared to AdamW in the pre-training of Llama models (of sizes ranging from 60M to 229M) on datasets including Wikitext-103, Minipile, and Openwebtext. Moreover, we provide theoretical guarantees, showing that IRE can substantially accelerate the convergence towards flat minima in Sharpness-aware Minimization (SAM).
title Improving Generalization and Convergence by Enhancing Implicit Regularization
topic Machine Learning
Optimization and Control
url https://arxiv.org/abs/2405.20763