Characterizing the Generalization Error of Random Feature Regression with Arbitrary Data-Augmentation
Fuente:
arXiv
Salvato in:
| Autori principali: | , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866916000267501568 |
|---|---|
| author | Morisset, Lucas Durmus, Alain Hardy, Adrien |
| author_facet | Morisset, Lucas Durmus, Alain Hardy, Adrien |
| contents | This paper aims at analyzing the regularization effect that data augmentation induces on supervised regression methods in the proportional regime, where the number of covariates grows proportionally to the number of samples. We provide a tight characterization of the test error, measured in mean squared error, in terms only of the population quantities of the true data, as well as first and second order statistics of the augmentation scheme. Our results are valid under misspecified feature maps, and for any network architecture where only the last readout layer is trained, and the rest of the network is either frozen or randomly initialized. We specify our results in the case of Gaussian data, and show that our asymptotic characterization is tight in this setting. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2605_10290 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Characterizing the Generalization Error of Random Feature Regression with Arbitrary Data-Augmentation Morisset, Lucas Durmus, Alain Hardy, Adrien Machine Learning Statistics Theory This paper aims at analyzing the regularization effect that data augmentation induces on supervised regression methods in the proportional regime, where the number of covariates grows proportionally to the number of samples. We provide a tight characterization of the test error, measured in mean squared error, in terms only of the population quantities of the true data, as well as first and second order statistics of the augmentation scheme. Our results are valid under misspecified feature maps, and for any network architecture where only the last readout layer is trained, and the rest of the network is either frozen or randomly initialized. We specify our results in the case of Gaussian data, and show that our asymptotic characterization is tight in this setting. |
| title | Characterizing the Generalization Error of Random Feature Regression with Arbitrary Data-Augmentation |
| topic | Machine Learning Statistics Theory |
| url | https://arxiv.org/abs/2605.10290 |