On identification in ill-posed linear regression

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Finocchio, Gianluca, Krivobokova, Tatyana
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914368321486848
author Finocchio, Gianluca
Krivobokova, Tatyana
author_facet Finocchio, Gianluca
Krivobokova, Tatyana
contents A novel framework is introduced to formalize identifiability in well-specified but ill-posed linear regression models. The framework is distribution-free and accommodates highly correlated features that may or may not relate to the response, reflecting typical real-data structures. First, the identifiable parameter is defined as the least-squares solution obtained by regressing the response on the largest subset of relevant features whose condition number does not exceed a specified threshold, and the relative risk incurred by using this predictor instead of the optimal one is quantified. Second, simple, verifiable conditions are provided under which a broad class of linear dimensionality reduction algorithms can estimate identifiable parameters; algorithms satisfying these conditions are termed statistically interpretable. Third, sharp high-probability error bounds are derived for these algorithms, with rates explicitly reflecting the degree of ill-posedness. With heavy-tailed features and sufficiently low effective rank, these algorithms achieve convergence rates that improve upon both the minimax least-squares rate and lower bounds for sparse estimation under sub-Gaussian features. Results are illustrated via simulations and a real-data application, in which effective rank grows logarithmically with dimension. The framework may extend to algorithms modeling nonlinear response-feature dependence.
format Preprint
id arxiv_https___arxiv_org_abs_2505_01297
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle On identification in ill-posed linear regression
Finocchio, Gianluca
Krivobokova, Tatyana
Statistics Theory
65F22, 65F10 (Primary) 62B05, 65F20 (Secondary)
A novel framework is introduced to formalize identifiability in well-specified but ill-posed linear regression models. The framework is distribution-free and accommodates highly correlated features that may or may not relate to the response, reflecting typical real-data structures. First, the identifiable parameter is defined as the least-squares solution obtained by regressing the response on the largest subset of relevant features whose condition number does not exceed a specified threshold, and the relative risk incurred by using this predictor instead of the optimal one is quantified. Second, simple, verifiable conditions are provided under which a broad class of linear dimensionality reduction algorithms can estimate identifiable parameters; algorithms satisfying these conditions are termed statistically interpretable. Third, sharp high-probability error bounds are derived for these algorithms, with rates explicitly reflecting the degree of ill-posedness. With heavy-tailed features and sufficiently low effective rank, these algorithms achieve convergence rates that improve upon both the minimax least-squares rate and lower bounds for sparse estimation under sub-Gaussian features. Results are illustrated via simulations and a real-data application, in which effective rank grows logarithmically with dimension. The framework may extend to algorithms modeling nonlinear response-feature dependence.
title On identification in ill-posed linear regression
topic Statistics Theory
65F22, 65F10 (Primary) 62B05, 65F20 (Secondary)
url https://arxiv.org/abs/2505.01297