A Unified Evaluation Framework for Epistemic Predictions

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Manchingal, Shireen Kudukkil, Mubashar, Muhammad, Wang, Kaizheng, Cuzzolin, Fabio
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929715684573184
author Manchingal, Shireen Kudukkil
Mubashar, Muhammad
Wang, Kaizheng
Cuzzolin, Fabio
author_facet Manchingal, Shireen Kudukkil
Mubashar, Muhammad
Wang, Kaizheng
Cuzzolin, Fabio
contents Predictions of uncertainty-aware models are diverse, ranging from single point estimates (often averaged over prediction samples) to predictive distributions, to set-valued or credal-set representations. We propose a novel unified evaluation framework for uncertainty-aware classifiers, applicable to a wide range of model classes, which allows users to tailor the trade-off between accuracy and precision of predictions via a suitably designed performance metric. This makes possible the selection of the most suitable model for a particular real-world application as a function of the desired trade-off. Our experiments, concerning Bayesian, ensemble, evidential, deterministic, credal and belief function classifiers on the CIFAR-10, MNIST and CIFAR-100 datasets, show that the metric behaves as desired.
format Preprint
id arxiv_https___arxiv_org_abs_2501_16912
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Unified Evaluation Framework for Epistemic Predictions
Manchingal, Shireen Kudukkil
Mubashar, Muhammad
Wang, Kaizheng
Cuzzolin, Fabio
Machine Learning
Predictions of uncertainty-aware models are diverse, ranging from single point estimates (often averaged over prediction samples) to predictive distributions, to set-valued or credal-set representations. We propose a novel unified evaluation framework for uncertainty-aware classifiers, applicable to a wide range of model classes, which allows users to tailor the trade-off between accuracy and precision of predictions via a suitably designed performance metric. This makes possible the selection of the most suitable model for a particular real-world application as a function of the desired trade-off. Our experiments, concerning Bayesian, ensemble, evidential, deterministic, credal and belief function classifiers on the CIFAR-10, MNIST and CIFAR-100 datasets, show that the metric behaves as desired.
title A Unified Evaluation Framework for Epistemic Predictions
topic Machine Learning
url https://arxiv.org/abs/2501.16912