All Models Are Miscalibrated, But Some Less So: Comparing Calibration with Conditional Mean Operators

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Moskvichev, Peter, Sejdinovic, Dino
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917149611655168
author Moskvichev, Peter
Sejdinovic, Dino
author_facet Moskvichev, Peter
Sejdinovic, Dino
contents When working in a high-risk setting, having well calibrated probabilistic predictive models is a crucial requirement. However, estimators for calibration error are not always able to correctly distinguish which model is better calibrated. We propose the \emph{conditional kernel calibration error} (CKCE) which is based on the Hilbert-Schmidt norm of the difference between conditional mean operators. By working directly with the definition of strong calibration as the distance between conditional distributions, which we represent by their embeddings in reproducing kernel Hilbert spaces, the CKCE is less sensitive to the marginal distribution of predictive models. This makes it more effective for relative comparisons than previously proposed calibration metrics. Our experiments, using both synthetic and real data, show that CKCE provides a more consistent ranking of models by their calibration error and is more robust against distribution shift.
format Preprint
id arxiv_https___arxiv_org_abs_2502_11465
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle All Models Are Miscalibrated, But Some Less So: Comparing Calibration with Conditional Mean Operators
Moskvichev, Peter
Sejdinovic, Dino
Machine Learning
When working in a high-risk setting, having well calibrated probabilistic predictive models is a crucial requirement. However, estimators for calibration error are not always able to correctly distinguish which model is better calibrated. We propose the \emph{conditional kernel calibration error} (CKCE) which is based on the Hilbert-Schmidt norm of the difference between conditional mean operators. By working directly with the definition of strong calibration as the distance between conditional distributions, which we represent by their embeddings in reproducing kernel Hilbert spaces, the CKCE is less sensitive to the marginal distribution of predictive models. This makes it more effective for relative comparisons than previously proposed calibration metrics. Our experiments, using both synthetic and real data, show that CKCE provides a more consistent ranking of models by their calibration error and is more robust against distribution shift.
title All Models Are Miscalibrated, But Some Less So: Comparing Calibration with Conditional Mean Operators
topic Machine Learning
url https://arxiv.org/abs/2502.11465