On Regularization via Early Stopping for Least Squares Regression

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Sonthalia, Rishi, Lok, Jackie, Rebrova, Elizaveta
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909218652553216
author Sonthalia, Rishi
Lok, Jackie
Rebrova, Elizaveta
author_facet Sonthalia, Rishi
Lok, Jackie
Rebrova, Elizaveta
contents A fundamental problem in machine learning is understanding the effect of early stopping on the parameters obtained and the generalization capabilities of the model. Even for linear models, the effect is not fully understood for arbitrary learning rates and data. In this paper, we analyze the dynamics of discrete full batch gradient descent for linear regression. With minimal assumptions, we characterize the trajectory of the parameters and the expected excess risk. Using this characterization, we show that when training with a learning rate schedule $η_k$, and a finite time horizon $T$, the early stopped solution $β_T$ is equivalent to the minimum norm solution for a generalized ridge regularized problem. We also prove that early stopping is beneficial for generic data with arbitrary spectrum and for a wide variety of learning rate schedules. We provide an estimate for the optimal stopping time and empirically demonstrate the accuracy of our estimate.
format Preprint
id arxiv_https___arxiv_org_abs_2406_04425
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle On Regularization via Early Stopping for Least Squares Regression
Sonthalia, Rishi
Lok, Jackie
Rebrova, Elizaveta
Machine Learning
Optimization and Control
Statistics Theory
A fundamental problem in machine learning is understanding the effect of early stopping on the parameters obtained and the generalization capabilities of the model. Even for linear models, the effect is not fully understood for arbitrary learning rates and data. In this paper, we analyze the dynamics of discrete full batch gradient descent for linear regression. With minimal assumptions, we characterize the trajectory of the parameters and the expected excess risk. Using this characterization, we show that when training with a learning rate schedule $η_k$, and a finite time horizon $T$, the early stopped solution $β_T$ is equivalent to the minimum norm solution for a generalized ridge regularized problem. We also prove that early stopping is beneficial for generic data with arbitrary spectrum and for a wide variety of learning rate schedules. We provide an estimate for the optimal stopping time and empirically demonstrate the accuracy of our estimate.
title On Regularization via Early Stopping for Least Squares Regression
topic Machine Learning
Optimization and Control
Statistics Theory
url https://arxiv.org/abs/2406.04425