From Tempered to Benign Overfitting in ReLU Neural Networks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kornowski, Guy, Yehudai, Gilad, Shamir, Ohad
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909143970873344
author Kornowski, Guy
Yehudai, Gilad
Shamir, Ohad
author_facet Kornowski, Guy
Yehudai, Gilad
Shamir, Ohad
contents Overparameterized neural networks (NNs) are observed to generalize well even when trained to perfectly fit noisy data. This phenomenon motivated a large body of work on "benign overfitting", where interpolating predictors achieve near-optimal performance. Recently, it was conjectured and empirically observed that the behavior of NNs is often better described as "tempered overfitting", where the performance is non-optimal yet also non-trivial, and degrades as a function of the noise level. However, a theoretical justification of this claim for non-linear NNs has been lacking so far. In this work, we provide several results that aim at bridging these complementing views. We study a simple classification setting with 2-layer ReLU NNs, and prove that under various assumptions, the type of overfitting transitions from tempered in the extreme case of one-dimensional data, to benign in high dimensions. Thus, we show that the input dimension has a crucial role on the type of overfitting in this setting, which we also validate empirically for intermediate dimensions. Overall, our results shed light on the intricate connections between the dimension, sample size, architecture and training algorithm on the one hand, and the type of resulting overfitting on the other hand.
format Preprint
id arxiv_https___arxiv_org_abs_2305_15141
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle From Tempered to Benign Overfitting in ReLU Neural Networks
Kornowski, Guy
Yehudai, Gilad
Shamir, Ohad
Machine Learning
Neural and Evolutionary Computing
Overparameterized neural networks (NNs) are observed to generalize well even when trained to perfectly fit noisy data. This phenomenon motivated a large body of work on "benign overfitting", where interpolating predictors achieve near-optimal performance. Recently, it was conjectured and empirically observed that the behavior of NNs is often better described as "tempered overfitting", where the performance is non-optimal yet also non-trivial, and degrades as a function of the noise level. However, a theoretical justification of this claim for non-linear NNs has been lacking so far. In this work, we provide several results that aim at bridging these complementing views. We study a simple classification setting with 2-layer ReLU NNs, and prove that under various assumptions, the type of overfitting transitions from tempered in the extreme case of one-dimensional data, to benign in high dimensions. Thus, we show that the input dimension has a crucial role on the type of overfitting in this setting, which we also validate empirically for intermediate dimensions. Overall, our results shed light on the intricate connections between the dimension, sample size, architecture and training algorithm on the one hand, and the type of resulting overfitting on the other hand.
title From Tempered to Benign Overfitting in ReLU Neural Networks
topic Machine Learning
Neural and Evolutionary Computing
url https://arxiv.org/abs/2305.15141