Confidence intervals uncovered: Are we ready for real-world medical imaging AI?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Christodoulou, Evangelia, Reinke, Annika, Houhou, Rola, Kalinowski, Piotr, Erkan, Selen, Sudre, Carole H., Burgos, Ninon, Boutaj, Sofiène, Loizillon, Sophie, Solal, Maëlys, Rieke, Nicola, Cheplygina, Veronika, Antonelli, Michela, Mayer, Leon D., Tizabi, Minu D., Cardoso, M. Jorge, Simpson, Amber, Jäger, Paul F., Kopp-Schneider, Annette, Varoquaux, Gaël, Colliot, Olivier, Maier-Hein, Lena
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913520655794176
author Christodoulou, Evangelia
Reinke, Annika
Houhou, Rola
Kalinowski, Piotr
Erkan, Selen
Sudre, Carole H.
Burgos, Ninon
Boutaj, Sofiène
Loizillon, Sophie
Solal, Maëlys
Rieke, Nicola
Cheplygina, Veronika
Antonelli, Michela
Mayer, Leon D.
Tizabi, Minu D.
Cardoso, M. Jorge
Simpson, Amber
Jäger, Paul F.
Kopp-Schneider, Annette
Varoquaux, Gaël
Colliot, Olivier
Maier-Hein, Lena
author_facet Christodoulou, Evangelia
Reinke, Annika
Houhou, Rola
Kalinowski, Piotr
Erkan, Selen
Sudre, Carole H.
Burgos, Ninon
Boutaj, Sofiène
Loizillon, Sophie
Solal, Maëlys
Rieke, Nicola
Cheplygina, Veronika
Antonelli, Michela
Mayer, Leon D.
Tizabi, Minu D.
Cardoso, M. Jorge
Simpson, Amber
Jäger, Paul F.
Kopp-Schneider, Annette
Varoquaux, Gaël
Colliot, Olivier
Maier-Hein, Lena
contents Medical imaging is spearheading the AI transformation of healthcare. Performance reporting is key to determine which methods should be translated into clinical practice. Frequently, broad conclusions are simply derived from mean performance values. In this paper, we argue that this common practice is often a misleading simplification as it ignores performance variability. Our contribution is threefold. (1) Analyzing all MICCAI segmentation papers (n = 221) published in 2023, we first observe that more than 50% of papers do not assess performance variability at all. Moreover, only one (0.5%) paper reported confidence intervals (CIs) for model performance. (2) To address the reporting bottleneck, we show that the unreported standard deviation (SD) in segmentation papers can be approximated by a second-order polynomial function of the mean Dice similarity coefficient (DSC). Based on external validation data from 56 previous MICCAI challenges, we demonstrate that this approximation can accurately reconstruct the CI of a method using information provided in publications. (3) Finally, we reconstructed 95% CIs around the mean DSC of MICCAI 2023 segmentation papers. The median CI width was 0.03 which is three times larger than the median performance gap between the first and second ranked method. For more than 60% of papers, the mean performance of the second-ranked method was within the CI of the first-ranked method. We conclude that current publications typically do not provide sufficient evidence to support which models could potentially be translated into clinical practice.
format Preprint
id arxiv_https___arxiv_org_abs_2409_17763
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Confidence intervals uncovered: Are we ready for real-world medical imaging AI?
Christodoulou, Evangelia
Reinke, Annika
Houhou, Rola
Kalinowski, Piotr
Erkan, Selen
Sudre, Carole H.
Burgos, Ninon
Boutaj, Sofiène
Loizillon, Sophie
Solal, Maëlys
Rieke, Nicola
Cheplygina, Veronika
Antonelli, Michela
Mayer, Leon D.
Tizabi, Minu D.
Cardoso, M. Jorge
Simpson, Amber
Jäger, Paul F.
Kopp-Schneider, Annette
Varoquaux, Gaël
Colliot, Olivier
Maier-Hein, Lena
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Medical imaging is spearheading the AI transformation of healthcare. Performance reporting is key to determine which methods should be translated into clinical practice. Frequently, broad conclusions are simply derived from mean performance values. In this paper, we argue that this common practice is often a misleading simplification as it ignores performance variability. Our contribution is threefold. (1) Analyzing all MICCAI segmentation papers (n = 221) published in 2023, we first observe that more than 50% of papers do not assess performance variability at all. Moreover, only one (0.5%) paper reported confidence intervals (CIs) for model performance. (2) To address the reporting bottleneck, we show that the unreported standard deviation (SD) in segmentation papers can be approximated by a second-order polynomial function of the mean Dice similarity coefficient (DSC). Based on external validation data from 56 previous MICCAI challenges, we demonstrate that this approximation can accurately reconstruct the CI of a method using information provided in publications. (3) Finally, we reconstructed 95% CIs around the mean DSC of MICCAI 2023 segmentation papers. The median CI width was 0.03 which is three times larger than the median performance gap between the first and second ranked method. For more than 60% of papers, the mean performance of the second-ranked method was within the CI of the first-ranked method. We conclude that current publications typically do not provide sufficient evidence to support which models could potentially be translated into clinical practice.
title Confidence intervals uncovered: Are we ready for real-world medical imaging AI?
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2409.17763