Performance uncertainty in medical image analysis: a large-scale investigation of confidence intervals
Fuente:
arXiv
Saved in:
| Main Authors: | André, Pascaline, Heitz, Charles, Christodoulou, Evangelia, Reinke, Annika, Sudre, Carole H., Antonelli, Michela, Godau, Patrick, Cardoso, M. Jorge, Gilson, Antoine, Montcel, Sophie Tezenas du, Varoquaux, Gaël, Maier-Hein, Lena, Colliot, Olivier |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Confidence intervals uncovered: Are we ready for real-world medical imaging AI?
by: Christodoulou, Evangelia, et al.
Published: (2024)
by: Christodoulou, Evangelia, et al.
Published: (2024)
False Promises in Medical Imaging AI? Assessing Validity of Outperformance Claims
by: Christodoulou, Evangelia, et al.
Published: (2025)
by: Christodoulou, Evangelia, et al.
Published: (2025)
A joint spatiotemporal model for multiple longitudinal markers and competing events
by: Ortholand, Juliette, et al.
Published: (2025)
by: Ortholand, Juliette, et al.
Published: (2025)
Joint model with latent disease age: overcoming the need for reference time
by: Ortholand, Juliette, et al.
Published: (2024)
by: Ortholand, Juliette, et al.
Published: (2024)
Challenging Vision-Language Models with Surgical Data: A New Dataset and Broad Benchmarking Study
by: Mayer, Leon, et al.
Published: (2025)
by: Mayer, Leon, et al.
Published: (2025)
A mixture model for subtype identification in the context of disease progression modeling
by: Kaisaridi, Sofia, et al.
Published: (2026)
by: Kaisaridi, Sofia, et al.
Published: (2026)
Quantifying Placebo Effects in Hereditary Ataxia Trials: A Meta‐Analysis of Scale for the Assessment and Rating of Ataxia ( SARA) Score Changes
by: Emilien Petit, et al.
Published: (2026)
by: Emilien Petit, et al.
Published: (2026)
Exploring Neural Correlates of Cognitive Awareness across the Alzheimer’s Disease Continuum: A Multimodal Study
by: Federica Cacciamani, et al.
Published: (2024)
by: Federica Cacciamani, et al.
Published: (2024)
Advancing standards in biomedical image analysis validation: A perspective on Metrics Reloaded
by: Annika Reinke, et al.
Published: (2025)
by: Annika Reinke, et al.
Published: (2025)
6 Fingers, 1 Kidney: Natural Adversarial Medical Images Reveal Critical Weaknesses of Vision-Language Models
by: Mayer, Leon, et al.
Published: (2025)
by: Mayer, Leon, et al.
Published: (2025)
Confidence Intervals for Performance Estimates in Brain MRI Segmentation
by: Jurdi, R. El, et al.
Published: (2023)
by: Jurdi, R. El, et al.
Published: (2023)
Beyond Knowledge Silos: Task Fingerprinting for Democratization of Medical Imaging AI
by: Godau, Patrick, et al.
Published: (2024)
by: Godau, Patrick, et al.
Published: (2024)
What is the Role of Small Models in the LLM Era: A Survey
by: Chen, Lihu, et al.
Published: (2024)
by: Chen, Lihu, et al.
Published: (2024)
Scalable Feature Learning on Huge Knowledge Graphs for Downstream Machine Learning
by: Lefebvre, Félix, et al.
Published: (2025)
by: Lefebvre, Félix, et al.
Published: (2025)
Imputation for prediction: beware of diminishing returns
by: Morvan, Marine Le, et al.
Published: (2024)
by: Morvan, Marine Le, et al.
Published: (2024)
Incidence, mortality, and management of status epilepticus from 2012 to 2022: An 11‐year nationwide study
by: Quentin Calonge, et al.
Published: (2025)
by: Quentin Calonge, et al.
Published: (2025)
Quantifying the severity of white matter damage in Cerebral Small Vessel Disease
by: Stylianos Charalampous, et al.
Published: (2025)
by: Stylianos Charalampous, et al.
Published: (2025)
Bond strength uncertainty quantification via confidence intervals for nondestructive evaluation of bonded composites
by: Stanley, Michael C., et al.
Published: (2025)
by: Stanley, Michael C., et al.
Published: (2025)
Learning High-Quality and General-Purpose Phrase Representations
by: Chen, Lihu, et al.
Published: (2024)
by: Chen, Lihu, et al.
Published: (2024)
CARTE: Pretraining and Transfer for Tabular Learning
by: Kim, Myung Jun, et al.
Published: (2024)
by: Kim, Myung Jun, et al.
Published: (2024)
Hype, Sustainability, and the Price of the Bigger-is-Better Paradigm in AI
by: Varoquaux, Gaël, et al.
Published: (2024)
by: Varoquaux, Gaël, et al.
Published: (2024)
On the calibration of survival models with competing risks
by: Alberge, Julie, et al.
Published: (2026)
by: Alberge, Julie, et al.
Published: (2026)
Risk ratio, odds ratio, risk difference... Which causal measure is easier to generalize?
by: Colnet, Bénédicte, et al.
Published: (2023)
by: Colnet, Bénédicte, et al.
Published: (2023)
Reweighting the RCT for generalization: finite sample error and variable selection
by: Colnet, Bénédicte, et al.
Published: (2022)
by: Colnet, Bénédicte, et al.
Published: (2022)
Bridging vision language model (VLM) evaluation gaps with a framework for scalable and cost-effective benchmark generation
by: Rädsch, Tim, et al.
Published: (2025)
by: Rädsch, Tim, et al.
Published: (2025)
Better bootstrap t confidence intervals for the mean
by: Owen, Art B.
Published: (2025)
by: Owen, Art B.
Published: (2025)
Optimal confidence interval for the difference of proportions
by: Peer, Almog, et al.
Published: (2023)
by: Peer, Almog, et al.
Published: (2023)
Fig.13. The average confidence interval
by: Al-Sammarraie, Dina, et al.
Published: (2020)
by: Al-Sammarraie, Dina, et al.
Published: (2020)
Fig. 12. The average confidence interval
by: Al-Sammarraie, Dina, et al.
Published: (2020)
by: Al-Sammarraie, Dina, et al.
Published: (2020)
Quality Assured: Rethinking Annotation Strategies in Imaging AI
by: Rädsch, Tim, et al.
Published: (2024)
by: Rädsch, Tim, et al.
Published: (2024)
Large Language Models as Search Engines: Societal Challenges
by: Sadeddine, Zacchary, et al.
Published: (2025)
by: Sadeddine, Zacchary, et al.
Published: (2025)
TabICL: A Tabular Foundation Model for In-Context Learning on Large Data
by: Qu, Jingang, et al.
Published: (2025)
by: Qu, Jingang, et al.
Published: (2025)
TabICLv2: A better, faster, scalable, and open tabular foundation model
by: Qu, Jingang, et al.
Published: (2026)
by: Qu, Jingang, et al.
Published: (2026)
Asymptotic optimality theory of confidence intervals of the mean
by: Deep, Vikas, et al.
Published: (2025)
by: Deep, Vikas, et al.
Published: (2025)
Likelihood confidence intervals for misspecified Cox models
by: Shao, Yongwu, et al.
Published: (2025)
by: Shao, Yongwu, et al.
Published: (2025)
Robust confidence intervals for generalized linear models
by: Panarotto, Andrea, et al.
Published: (2026)
by: Panarotto, Andrea, et al.
Published: (2026)
Tighter confidence intervals for quantiles of heterogeneous data
by: Einmahl, John H. J., et al.
Published: (2026)
by: Einmahl, John H. J., et al.
Published: (2026)
Robust single-stage selection problems with budgeted interval uncertainty
by: Lhomme, Antoine, et al.
Published: (2025)
by: Lhomme, Antoine, et al.
Published: (2025)
O conceito de política posto à prova pela mundialização
by: Catherine Colliot-Thélène
Published: (1999)
by: Catherine Colliot-Thélène
Published: (1999)
FISBe: A real-world benchmark dataset for instance segmentation of long-range thin filamentous structures
by: Mais, Lisa, et al.
Published: (2024)
by: Mais, Lisa, et al.
Published: (2024)
Similar Items
-
Confidence intervals uncovered: Are we ready for real-world medical imaging AI?
by: Christodoulou, Evangelia, et al.
Published: (2024) -
False Promises in Medical Imaging AI? Assessing Validity of Outperformance Claims
by: Christodoulou, Evangelia, et al.
Published: (2025) -
A joint spatiotemporal model for multiple longitudinal markers and competing events
by: Ortholand, Juliette, et al.
Published: (2025) -
Joint model with latent disease age: overcoming the need for reference time
by: Ortholand, Juliette, et al.
Published: (2024) -
Challenging Vision-Language Models with Surgical Data: A New Dataset and Broad Benchmarking Study
by: Mayer, Leon, et al.
Published: (2025)