Model Evaluation in the Dark: Robust Classifier Metrics with Missing Labels

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Dervovic, Danial, Cashmore, Michael
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910919007666176
author Dervovic, Danial
Cashmore, Michael
author_facet Dervovic, Danial
Cashmore, Michael
contents Missing data in supervised learning is well-studied, but the specific issue of missing labels during model evaluation has been overlooked. Ignoring samples with missing values, a common solution, can introduce bias, especially when data is Missing Not At Random (MNAR). We propose a multiple imputation technique for evaluating classifiers using metrics such as precision, recall, and ROC-AUC. This method not only offers point estimates but also a predictive distribution for these quantities when labels are missing. We empirically show that the predictive distribution's location and shape are generally correct, even in the MNAR regime. Moreover, we establish that this distribution is approximately Gaussian and provide finite-sample convergence bounds. Additionally, a robustness proof is presented, confirming the validity of the approximation under a realistic error model.
format Preprint
id arxiv_https___arxiv_org_abs_2504_18385
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Model Evaluation in the Dark: Robust Classifier Metrics with Missing Labels
Dervovic, Danial
Cashmore, Michael
Machine Learning
Missing data in supervised learning is well-studied, but the specific issue of missing labels during model evaluation has been overlooked. Ignoring samples with missing values, a common solution, can introduce bias, especially when data is Missing Not At Random (MNAR). We propose a multiple imputation technique for evaluating classifiers using metrics such as precision, recall, and ROC-AUC. This method not only offers point estimates but also a predictive distribution for these quantities when labels are missing. We empirically show that the predictive distribution's location and shape are generally correct, even in the MNAR regime. Moreover, we establish that this distribution is approximately Gaussian and provide finite-sample convergence bounds. Additionally, a robustness proof is presented, confirming the validity of the approximation under a realistic error model.
title Model Evaluation in the Dark: Robust Classifier Metrics with Missing Labels
topic Machine Learning
url https://arxiv.org/abs/2504.18385