Optimal Cross-Validation for Sparse Linear Regression

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Cory-Wright, Ryan, Gómez, Andrés
Format: Preprint
Veröffentlicht: 2023
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912898172846080
author Cory-Wright, Ryan
Gómez, Andrés
author_facet Cory-Wright, Ryan
Gómez, Andrés
contents Given a high-dimensional covariate matrix and a response vector, ridge-regularized sparse linear regression selects a subset of features that explains the relationship between covariates and the response in an interpretable manner. To choose hyperparameters that control the sparsity level and amount of regularization, practitioners commonly use k-fold cross-validation. However, cross-validation substantially increases the computational cost of sparse regression as it requires solving many mixed-integer optimization problems (MIOs) for each hyperparameter combination. To address this computational burden, we derive computationally tractable relaxations of the k-fold cross-validation loss, facilitating hyperparameter selection while solving $50$--$80\%$ fewer MIOs in practice. Our computational results demonstrate, across eleven real-world UCI datasets, that exact MIO-based cross-validation can be competitive with mature software packages such as glmnet and L0Learn -particularly when the sample-to-feature ratio is small.
format Preprint
id arxiv_https___arxiv_org_abs_2306_14851
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Optimal Cross-Validation for Sparse Linear Regression
Cory-Wright, Ryan
Gómez, Andrés
Optimization and Control
Machine Learning
Methodology
Given a high-dimensional covariate matrix and a response vector, ridge-regularized sparse linear regression selects a subset of features that explains the relationship between covariates and the response in an interpretable manner. To choose hyperparameters that control the sparsity level and amount of regularization, practitioners commonly use k-fold cross-validation. However, cross-validation substantially increases the computational cost of sparse regression as it requires solving many mixed-integer optimization problems (MIOs) for each hyperparameter combination. To address this computational burden, we derive computationally tractable relaxations of the k-fold cross-validation loss, facilitating hyperparameter selection while solving $50$--$80\%$ fewer MIOs in practice. Our computational results demonstrate, across eleven real-world UCI datasets, that exact MIO-based cross-validation can be competitive with mature software packages such as glmnet and L0Learn -particularly when the sample-to-feature ratio is small.
title Optimal Cross-Validation for Sparse Linear Regression
topic Optimization and Control
Machine Learning
Methodology
url https://arxiv.org/abs/2306.14851