Mind the spikes: Benign overfitting of kernels and neural networks in fixed dimension

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Haas, Moritz, Holzmüller, David, von Luxburg, Ulrike, Steinwart, Ingo
Formato: Preprint
Publicado: 2023
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866913572210081792
author Haas, Moritz
Holzmüller, David
von Luxburg, Ulrike
Steinwart, Ingo
author_facet Haas, Moritz
Holzmüller, David
von Luxburg, Ulrike
Steinwart, Ingo
contents The success of over-parameterized neural networks trained to near-zero training error has caused great interest in the phenomenon of benign overfitting, where estimators are statistically consistent even though they interpolate noisy training data. While benign overfitting in fixed dimension has been established for some learning methods, current literature suggests that for regression with typical kernel methods and wide neural networks, benign overfitting requires a high-dimensional setting where the dimension grows with the sample size. In this paper, we show that the smoothness of the estimators, and not the dimension, is the key: benign overfitting is possible if and only if the estimator's derivatives are large enough. We generalize existing inconsistency results to non-interpolating models and more kernels to show that benign overfitting with moderate derivatives is impossible in fixed dimension. Conversely, we show that rate-optimal benign overfitting is possible for regression with a sequence of spiky-smooth kernels with large derivatives. Using neural tangent kernels, we translate our results to wide neural networks. We prove that while infinite-width networks do not overfit benignly with the ReLU activation, this can be fixed by adding small high-frequency fluctuations to the activation function. Our experiments verify that such neural networks, while overfitting, can indeed generalize well even on low-dimensional data sets.
format Preprint
id arxiv_https___arxiv_org_abs_2305_14077
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Mind the spikes: Benign overfitting of kernels and neural networks in fixed dimension
Haas, Moritz
Holzmüller, David
von Luxburg, Ulrike
Steinwart, Ingo
Machine Learning
Statistics Theory
The success of over-parameterized neural networks trained to near-zero training error has caused great interest in the phenomenon of benign overfitting, where estimators are statistically consistent even though they interpolate noisy training data. While benign overfitting in fixed dimension has been established for some learning methods, current literature suggests that for regression with typical kernel methods and wide neural networks, benign overfitting requires a high-dimensional setting where the dimension grows with the sample size. In this paper, we show that the smoothness of the estimators, and not the dimension, is the key: benign overfitting is possible if and only if the estimator's derivatives are large enough. We generalize existing inconsistency results to non-interpolating models and more kernels to show that benign overfitting with moderate derivatives is impossible in fixed dimension. Conversely, we show that rate-optimal benign overfitting is possible for regression with a sequence of spiky-smooth kernels with large derivatives. Using neural tangent kernels, we translate our results to wide neural networks. We prove that while infinite-width networks do not overfit benignly with the ReLU activation, this can be fixed by adding small high-frequency fluctuations to the activation function. Our experiments verify that such neural networks, while overfitting, can indeed generalize well even on low-dimensional data sets.
title Mind the spikes: Benign overfitting of kernels and neural networks in fixed dimension
topic Machine Learning
Statistics Theory
url https://arxiv.org/abs/2305.14077