Batches Stabilize the Minimum Norm Risk in High Dimensional Overparameterized Linear Regression

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Ioushua, Shahar Stein, Hasidim, Inbar, Shayevitz, Ofer, Feder, Meir
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910615385145344
author Ioushua, Shahar Stein
Hasidim, Inbar
Shayevitz, Ofer
Feder, Meir
author_facet Ioushua, Shahar Stein
Hasidim, Inbar
Shayevitz, Ofer
Feder, Meir
contents Learning algorithms that divide the data into batches are prevalent in many machine-learning applications, typically offering useful trade-offs between computational efficiency and performance. In this paper, we examine the benefits of batch-partitioning through the lens of a minimum-norm overparametrized linear regression model with isotropic Gaussian features. We suggest a natural small-batch version of the minimum-norm estimator and derive bounds on its quadratic risk. We then characterize the optimal batch size and show it is inversely proportional to the noise level, as well as to the overparametrization ratio. In contrast to minimum-norm, our estimator admits a stable risk behavior that is monotonically increasing in the overparametrization ratio, eliminating both the blowup at the interpolation point and the double-descent phenomenon. We further show that shrinking the batch minimum-norm estimator by a factor equal to the Weiner coefficient further stabilizes it and results in lower quadratic risk in all settings. Interestingly, we observe that the implicit regularization offered by the batch partition is partially explained by feature overlap between the batches. Our bound is derived via a novel combination of techniques, in particular normal approximation in the Wasserstein metric of noisy projections over random subspaces.
format Preprint
id arxiv_https___arxiv_org_abs_2306_08432
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Batches Stabilize the Minimum Norm Risk in High Dimensional Overparameterized Linear Regression
Ioushua, Shahar Stein
Hasidim, Inbar
Shayevitz, Ofer
Feder, Meir
Machine Learning
Information Theory
Statistics Theory
Learning algorithms that divide the data into batches are prevalent in many machine-learning applications, typically offering useful trade-offs between computational efficiency and performance. In this paper, we examine the benefits of batch-partitioning through the lens of a minimum-norm overparametrized linear regression model with isotropic Gaussian features. We suggest a natural small-batch version of the minimum-norm estimator and derive bounds on its quadratic risk. We then characterize the optimal batch size and show it is inversely proportional to the noise level, as well as to the overparametrization ratio. In contrast to minimum-norm, our estimator admits a stable risk behavior that is monotonically increasing in the overparametrization ratio, eliminating both the blowup at the interpolation point and the double-descent phenomenon. We further show that shrinking the batch minimum-norm estimator by a factor equal to the Weiner coefficient further stabilizes it and results in lower quadratic risk in all settings. Interestingly, we observe that the implicit regularization offered by the batch partition is partially explained by feature overlap between the batches. Our bound is derived via a novel combination of techniques, in particular normal approximation in the Wasserstein metric of noisy projections over random subspaces.
title Batches Stabilize the Minimum Norm Risk in High Dimensional Overparameterized Linear Regression
topic Machine Learning
Information Theory
Statistics Theory
url https://arxiv.org/abs/2306.08432