A Honest Cross-Validation Estimator for Prediction Performance

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Pan, Tianyu, Yu, Vincent Z., Devanarayan, Viswanath, Tian, Lu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916997947719680
author Pan, Tianyu
Yu, Vincent Z.
Devanarayan, Viswanath
Tian, Lu
author_facet Pan, Tianyu
Yu, Vincent Z.
Devanarayan, Viswanath
Tian, Lu
contents Cross-validation is a standard tool for obtaining a honest assessment of the performance of a prediction model. The commonly used version repeatedly splits data, trains the prediction model on the training set, evaluates the model performance on the test set, and averages the model performance across different data splits. A well-known criticism is that such cross-validation procedure does not directly estimate the performance of the particular model recommended for future use. In this paper, we propose a new method to estimate the performance of a model trained on a specific (random) training set. A naive estimator can be obtained by applying the model to a disjoint testing set. Surprisingly, cross-validation estimators computed from other random splits can be used to improve this naive estimator within a random-effects model framework. We develop two estimators -- a hierarchical Bayesian estimator and an empirical Bayes estimator -- that perform similarly to or better than both the conventional cross-validation estimator and the naive single-split estimator. Simulations and a real-data example demonstrate the superior performance of the proposed method.
format Preprint
id arxiv_https___arxiv_org_abs_2510_07649
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Honest Cross-Validation Estimator for Prediction Performance
Pan, Tianyu
Yu, Vincent Z.
Devanarayan, Viswanath
Tian, Lu
Machine Learning
Applications
Methodology
Cross-validation is a standard tool for obtaining a honest assessment of the performance of a prediction model. The commonly used version repeatedly splits data, trains the prediction model on the training set, evaluates the model performance on the test set, and averages the model performance across different data splits. A well-known criticism is that such cross-validation procedure does not directly estimate the performance of the particular model recommended for future use. In this paper, we propose a new method to estimate the performance of a model trained on a specific (random) training set. A naive estimator can be obtained by applying the model to a disjoint testing set. Surprisingly, cross-validation estimators computed from other random splits can be used to improve this naive estimator within a random-effects model framework. We develop two estimators -- a hierarchical Bayesian estimator and an empirical Bayes estimator -- that perform similarly to or better than both the conventional cross-validation estimator and the naive single-split estimator. Simulations and a real-data example demonstrate the superior performance of the proposed method.
title A Honest Cross-Validation Estimator for Prediction Performance
topic Machine Learning
Applications
Methodology
url https://arxiv.org/abs/2510.07649