On Arbitrary Predictions from Equally Valid Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lockfisch, Sarah, Schwethelm, Kristian, Menten, Martin, Braren, Rickmer, Rueckert, Daniel, Ziller, Alexander, Kaissis, Georgios
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908903772520448
author Lockfisch, Sarah
Schwethelm, Kristian
Menten, Martin
Braren, Rickmer
Rueckert, Daniel
Ziller, Alexander
Kaissis, Georgios
author_facet Lockfisch, Sarah
Schwethelm, Kristian
Menten, Martin
Braren, Rickmer
Rueckert, Daniel
Ziller, Alexander
Kaissis, Georgios
contents Model multiplicity refers to the existence of multiple machine learning models that describe the data equally well but may produce different predictions on individual samples. In medicine, these models can admit conflicting predictions for the same patient -- a risk that is poorly understood and insufficiently addressed. In this study, we empirically analyze the extent, drivers, and ramifications of predictive multiplicity across diverse medical tasks and model architectures, and show that even small ensembles can mitigate/eliminate predictive multiplicity in practice. Our analysis reveals that (1) standard validation metrics fail to identify a uniquely optimal model and (2) a substantial amount of predictions hinges on arbitrary choices made during model development. Using multiple models instead of a single model reveals instances where predictions differ across equally plausible models -- highlighting patients that would receive arbitrary diagnoses if any single model were used. In contrast, (3) a small ensemble paired with an abstention strategy can effectively mitigate measurable predictive multiplicity in practice; predictions with high inter-model consensus may thus be amenable to automated classification. While accuracy is not a principled antidote to predictive multiplicity, we find that (4) higher accuracy achieved through increased model capacity reduces predictive multiplicity. Our findings underscore the clinical importance of accounting for model multiplicity and advocate for ensemble-based strategies to improve diagnostic reliability. In cases where models fail to reach sufficient consensus, we recommend deferring decisions to expert review.
format Preprint
id arxiv_https___arxiv_org_abs_2507_19408
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle On Arbitrary Predictions from Equally Valid Models
Lockfisch, Sarah
Schwethelm, Kristian
Menten, Martin
Braren, Rickmer
Rueckert, Daniel
Ziller, Alexander
Kaissis, Georgios
Machine Learning
Artificial Intelligence
Model multiplicity refers to the existence of multiple machine learning models that describe the data equally well but may produce different predictions on individual samples. In medicine, these models can admit conflicting predictions for the same patient -- a risk that is poorly understood and insufficiently addressed. In this study, we empirically analyze the extent, drivers, and ramifications of predictive multiplicity across diverse medical tasks and model architectures, and show that even small ensembles can mitigate/eliminate predictive multiplicity in practice. Our analysis reveals that (1) standard validation metrics fail to identify a uniquely optimal model and (2) a substantial amount of predictions hinges on arbitrary choices made during model development. Using multiple models instead of a single model reveals instances where predictions differ across equally plausible models -- highlighting patients that would receive arbitrary diagnoses if any single model were used. In contrast, (3) a small ensemble paired with an abstention strategy can effectively mitigate measurable predictive multiplicity in practice; predictions with high inter-model consensus may thus be amenable to automated classification. While accuracy is not a principled antidote to predictive multiplicity, we find that (4) higher accuracy achieved through increased model capacity reduces predictive multiplicity. Our findings underscore the clinical importance of accounting for model multiplicity and advocate for ensemble-based strategies to improve diagnostic reliability. In cases where models fail to reach sufficient consensus, we recommend deferring decisions to expert review.
title On Arbitrary Predictions from Equally Valid Models
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2507.19408