Label-free estimation of clinically relevant performance metrics under distribution shifts

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Flühmann, Tim, Bissoto, Alceu, Hoang, Trung-Dung, Koch, Lisa M.
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866916871301758976
author Flühmann, Tim
Bissoto, Alceu
Hoang, Trung-Dung
Koch, Lisa M.
author_facet Flühmann, Tim
Bissoto, Alceu
Hoang, Trung-Dung
Koch, Lisa M.
contents Performance monitoring is essential for safe clinical deployment of image classification models. However, because ground-truth labels are typically unavailable in the target dataset, direct assessment of real-world model performance is infeasible. State-of-the-art performance estimation methods address this by leveraging confidence scores to estimate the target accuracy. Despite being a promising direction, the established methods mainly estimate the model's accuracy and are rarely evaluated in a clinical domain, where strong class imbalances and dataset shifts are common. Our contributions are twofold: First, we introduce generalisations of existing performance prediction methods that directly estimate the full confusion matrix. Then, we benchmark their performance on chest x-ray data in real-world distribution shifts as well as simulated covariate and prevalence shifts. The proposed confusion matrix estimation methods reliably predicted clinically relevant counting metrics on medical images under distribution shifts. However, our simulated shift scenarios exposed important failure modes of current performance estimation techniques, calling for a better understanding of real-world deployment contexts when implementing these performance monitoring techniques for postmarket surveillance of medical AI models.
format Preprint
id arxiv_https___arxiv_org_abs_2507_22776
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Label-free estimation of clinically relevant performance metrics under distribution shifts
Flühmann, Tim
Bissoto, Alceu
Hoang, Trung-Dung
Koch, Lisa M.
Machine Learning
Performance monitoring is essential for safe clinical deployment of image classification models. However, because ground-truth labels are typically unavailable in the target dataset, direct assessment of real-world model performance is infeasible. State-of-the-art performance estimation methods address this by leveraging confidence scores to estimate the target accuracy. Despite being a promising direction, the established methods mainly estimate the model's accuracy and are rarely evaluated in a clinical domain, where strong class imbalances and dataset shifts are common. Our contributions are twofold: First, we introduce generalisations of existing performance prediction methods that directly estimate the full confusion matrix. Then, we benchmark their performance on chest x-ray data in real-world distribution shifts as well as simulated covariate and prevalence shifts. The proposed confusion matrix estimation methods reliably predicted clinically relevant counting metrics on medical images under distribution shifts. However, our simulated shift scenarios exposed important failure modes of current performance estimation techniques, calling for a better understanding of real-world deployment contexts when implementing these performance monitoring techniques for postmarket surveillance of medical AI models.
title Label-free estimation of clinically relevant performance metrics under distribution shifts
topic Machine Learning
url https://arxiv.org/abs/2507.22776