Understanding Model Calibration -- A gentle introduction and visual exploration of calibration and the expected calibration error (ECE)

Fuente: arXiv
Saved in:
Bibliographic Details
Main Author: Pavlovic, Maja
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908536257118208
author Pavlovic, Maja
author_facet Pavlovic, Maja
contents To be considered reliable, a model must be calibrated so that its confidence in each decision closely reflects its true outcome. In this blogpost we'll take a look at the most commonly used definition for calibration and then dive into a frequently used evaluation measure for model calibration. We'll then cover some of the drawbacks of this measure and how these surfaced the need for additional notions of calibration, which require their own new evaluation measures. This post is not intended to be an in-depth dissection of all works on calibration, nor does it focus on how to calibrate models. Instead, it is meant to provide a gentle introduction to the different notions and their evaluation measures as well as to re-highlight some issues with a measure that is still widely used to evaluate calibration.
format Preprint
id arxiv_https___arxiv_org_abs_2501_19047
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Understanding Model Calibration -- A gentle introduction and visual exploration of calibration and the expected calibration error (ECE)
Pavlovic, Maja
Methodology
Artificial Intelligence
Computer Vision and Pattern Recognition
Machine Learning
To be considered reliable, a model must be calibrated so that its confidence in each decision closely reflects its true outcome. In this blogpost we'll take a look at the most commonly used definition for calibration and then dive into a frequently used evaluation measure for model calibration. We'll then cover some of the drawbacks of this measure and how these surfaced the need for additional notions of calibration, which require their own new evaluation measures. This post is not intended to be an in-depth dissection of all works on calibration, nor does it focus on how to calibrate models. Instead, it is meant to provide a gentle introduction to the different notions and their evaluation measures as well as to re-highlight some issues with a measure that is still widely used to evaluate calibration.
title Understanding Model Calibration -- A gentle introduction and visual exploration of calibration and the expected calibration error (ECE)
topic Methodology
Artificial Intelligence
Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2501.19047