Calibration improves detection of mislabeled examples

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chibane, Ilies, George, Thomas, Nodet, Pierre, Lemaire, Vincent
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912687866249216
author Chibane, Ilies
George, Thomas
Nodet, Pierre
Lemaire, Vincent
author_facet Chibane, Ilies
George, Thomas
Nodet, Pierre
Lemaire, Vincent
contents Mislabeled data is a pervasive issue that undermines the performance of machine learning systems in real-world applications. An effective approach to mitigate this problem is to detect mislabeled instances and subject them to special treatment, such as filtering or relabeling. Automatic mislabeling detection methods typically rely on training a base machine learning model and then probing it for each instance to obtain a trust score that each provided label is genuine or incorrect. The properties of this base model are thus of paramount importance. In this paper, we investigate the impact of calibrating this model. Our empirical results show that using calibration methods improves the accuracy and robustness of mislabeled instance detection, providing a practical and effective solution for industrial applications.
format Preprint
id arxiv_https___arxiv_org_abs_2511_02738
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Calibration improves detection of mislabeled examples
Chibane, Ilies
George, Thomas
Nodet, Pierre
Lemaire, Vincent
Machine Learning
Mislabeled data is a pervasive issue that undermines the performance of machine learning systems in real-world applications. An effective approach to mitigate this problem is to detect mislabeled instances and subject them to special treatment, such as filtering or relabeling. Automatic mislabeling detection methods typically rely on training a base machine learning model and then probing it for each instance to obtain a trust score that each provided label is genuine or incorrect. The properties of this base model are thus of paramount importance. In this paper, we investigate the impact of calibrating this model. Our empirical results show that using calibration methods improves the accuracy and robustness of mislabeled instance detection, providing a practical and effective solution for industrial applications.
title Calibration improves detection of mislabeled examples
topic Machine Learning
url https://arxiv.org/abs/2511.02738