Characterizing the Generalization Error of Random Feature Regression with Arbitrary Data-Augmentation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Morisset, Lucas, Durmus, Alain, Hardy, Adrien
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916000267501568
author Morisset, Lucas
Durmus, Alain
Hardy, Adrien
author_facet Morisset, Lucas
Durmus, Alain
Hardy, Adrien
contents This paper aims at analyzing the regularization effect that data augmentation induces on supervised regression methods in the proportional regime, where the number of covariates grows proportionally to the number of samples. We provide a tight characterization of the test error, measured in mean squared error, in terms only of the population quantities of the true data, as well as first and second order statistics of the augmentation scheme. Our results are valid under misspecified feature maps, and for any network architecture where only the last readout layer is trained, and the rest of the network is either frozen or randomly initialized. We specify our results in the case of Gaussian data, and show that our asymptotic characterization is tight in this setting.
format Preprint
id arxiv_https___arxiv_org_abs_2605_10290
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Characterizing the Generalization Error of Random Feature Regression with Arbitrary Data-Augmentation
Morisset, Lucas
Durmus, Alain
Hardy, Adrien
Machine Learning
Statistics Theory
This paper aims at analyzing the regularization effect that data augmentation induces on supervised regression methods in the proportional regime, where the number of covariates grows proportionally to the number of samples. We provide a tight characterization of the test error, measured in mean squared error, in terms only of the population quantities of the true data, as well as first and second order statistics of the augmentation scheme. Our results are valid under misspecified feature maps, and for any network architecture where only the last readout layer is trained, and the rest of the network is either frozen or randomly initialized. We specify our results in the case of Gaussian data, and show that our asymptotic characterization is tight in this setting.
title Characterizing the Generalization Error of Random Feature Regression with Arbitrary Data-Augmentation
topic Machine Learning
Statistics Theory
url https://arxiv.org/abs/2605.10290