Optimal Data Splitting for Holdout Cross-Validation in Large Covariance Matrix Estimation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lamrani, Lamia, Bongiorno, Christian, Potters, Marc
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908544071106560
author Lamrani, Lamia
Bongiorno, Christian
Potters, Marc
author_facet Lamrani, Lamia
Bongiorno, Christian
Potters, Marc
contents Cross-validation is a statistical tool that can be used to improve large covariance matrix estimation. Although its efficiency is observed in practical applications and a convergence result towards the error of the non linear shrinkage is available in the high-dimensional regime, formal proofs that take into account the finite sample size effects are currently lacking. To carry on analytical analysis, we focus on the holdout method, a single iteration of cross-validation, rather than the traditional $k$-fold approach. We derive a closed-form expression for the expected estimation error when the population matrix follows a white inverse Wishart distribution, and we observe the optimal train-test split scales as the square root of the matrix dimension. For general population matrices, we connected the error to the variance of eigenvalues distribution, but approximations are necessary. In this framework and in the high-dimensional asymptotic regime, both the holdout and $k$-fold cross-validation methods converge to the optimal estimator when the train-test ratio scales with the square root of the matrix dimension which is coherent with the existing theory.
format Preprint
id arxiv_https___arxiv_org_abs_2503_15186
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Optimal Data Splitting for Holdout Cross-Validation in Large Covariance Matrix Estimation
Lamrani, Lamia
Bongiorno, Christian
Potters, Marc
Statistics Theory
Portfolio Management
Risk Management
Applications
Cross-validation is a statistical tool that can be used to improve large covariance matrix estimation. Although its efficiency is observed in practical applications and a convergence result towards the error of the non linear shrinkage is available in the high-dimensional regime, formal proofs that take into account the finite sample size effects are currently lacking. To carry on analytical analysis, we focus on the holdout method, a single iteration of cross-validation, rather than the traditional $k$-fold approach. We derive a closed-form expression for the expected estimation error when the population matrix follows a white inverse Wishart distribution, and we observe the optimal train-test split scales as the square root of the matrix dimension. For general population matrices, we connected the error to the variance of eigenvalues distribution, but approximations are necessary. In this framework and in the high-dimensional asymptotic regime, both the holdout and $k$-fold cross-validation methods converge to the optimal estimator when the train-test ratio scales with the square root of the matrix dimension which is coherent with the existing theory.
title Optimal Data Splitting for Holdout Cross-Validation in Large Covariance Matrix Estimation
topic Statistics Theory
Portfolio Management
Risk Management
Applications
url https://arxiv.org/abs/2503.15186