Noisy Interpolation Learning with Shallow Univariate ReLU Networks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Joshi, Nirmit, Vardi, Gal, Srebro, Nathan
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916171393007616
author Joshi, Nirmit
Vardi, Gal
Srebro, Nathan
author_facet Joshi, Nirmit
Vardi, Gal
Srebro, Nathan
contents Understanding how overparameterized neural networks generalize despite perfect interpolation of noisy training data is a fundamental question. Mallinar et. al. 2022 noted that neural networks seem to often exhibit ``tempered overfitting'', wherein the population risk does not converge to the Bayes optimal error, but neither does it approach infinity, yielding non-trivial generalization. However, this has not been studied rigorously. We provide the first rigorous analysis of the overfitting behavior of regression with minimum norm ($\ell_2$ of weights), focusing on univariate two-layer ReLU networks. We show overfitting is tempered (with high probability) when measured with respect to the $L_1$ loss, but also show that the situation is more complex than suggested by Mallinar et. al., and overfitting is catastrophic with respect to the $L_2$ loss, or when taking an expectation over the training set.
format Preprint
id arxiv_https___arxiv_org_abs_2307_15396
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Noisy Interpolation Learning with Shallow Univariate ReLU Networks
Joshi, Nirmit
Vardi, Gal
Srebro, Nathan
Machine Learning
Understanding how overparameterized neural networks generalize despite perfect interpolation of noisy training data is a fundamental question. Mallinar et. al. 2022 noted that neural networks seem to often exhibit ``tempered overfitting'', wherein the population risk does not converge to the Bayes optimal error, but neither does it approach infinity, yielding non-trivial generalization. However, this has not been studied rigorously. We provide the first rigorous analysis of the overfitting behavior of regression with minimum norm ($\ell_2$ of weights), focusing on univariate two-layer ReLU networks. We show overfitting is tempered (with high probability) when measured with respect to the $L_1$ loss, but also show that the situation is more complex than suggested by Mallinar et. al., and overfitting is catastrophic with respect to the $L_2$ loss, or when taking an expectation over the training set.
title Noisy Interpolation Learning with Shallow Univariate ReLU Networks
topic Machine Learning
url https://arxiv.org/abs/2307.15396