High-dimensional robust regression under heavy-tailed data: Asymptotics and Universality

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Adomaityte, Urte, Defilippis, Leonardo, Loureiro, Bruno, Sicuro, Gabriele
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914817460142080
author Adomaityte, Urte
Defilippis, Leonardo
Loureiro, Bruno
Sicuro, Gabriele
author_facet Adomaityte, Urte
Defilippis, Leonardo
Loureiro, Bruno
Sicuro, Gabriele
contents We investigate the high-dimensional properties of robust regression estimators in the presence of heavy-tailed contamination of both the covariates and response functions. In particular, we provide a sharp asymptotic characterisation of M-estimators trained on a family of elliptical covariate and noise data distributions including cases where second and higher moments do not exist. We show that, despite being consistent, the Huber loss with optimally tuned location parameter $δ$ is suboptimal in the high-dimensional regime in the presence of heavy-tailed noise, highlighting the necessity of further regularisation to achieve optimal performance. This result also uncovers the existence of a transition in $δ$ as a function of the sample complexity and contamination. Moreover, we derive the decay rates for the excess risk of ridge regression. We show that, while it is both optimal and universal for covariate distributions with finite second moment, its decay rate can be considerably faster when the covariates' second moment does not exist. Finally, we show that our formulas readily generalise to a richer family of models and data distributions, such as generalised linear estimation with arbitrary convex regularisation trained on mixture models.
format Preprint
id arxiv_https___arxiv_org_abs_2309_16476
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle High-dimensional robust regression under heavy-tailed data: Asymptotics and Universality
Adomaityte, Urte
Defilippis, Leonardo
Loureiro, Bruno
Sicuro, Gabriele
Statistics Theory
Disordered Systems and Neural Networks
Machine Learning
We investigate the high-dimensional properties of robust regression estimators in the presence of heavy-tailed contamination of both the covariates and response functions. In particular, we provide a sharp asymptotic characterisation of M-estimators trained on a family of elliptical covariate and noise data distributions including cases where second and higher moments do not exist. We show that, despite being consistent, the Huber loss with optimally tuned location parameter $δ$ is suboptimal in the high-dimensional regime in the presence of heavy-tailed noise, highlighting the necessity of further regularisation to achieve optimal performance. This result also uncovers the existence of a transition in $δ$ as a function of the sample complexity and contamination. Moreover, we derive the decay rates for the excess risk of ridge regression. We show that, while it is both optimal and universal for covariate distributions with finite second moment, its decay rate can be considerably faster when the covariates' second moment does not exist. Finally, we show that our formulas readily generalise to a richer family of models and data distributions, such as generalised linear estimation with arbitrary convex regularisation trained on mixture models.
title High-dimensional robust regression under heavy-tailed data: Asymptotics and Universality
topic Statistics Theory
Disordered Systems and Neural Networks
Machine Learning
url https://arxiv.org/abs/2309.16476