Detecting and Rectifying Noisy Labels: A Similarity-based Approach

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Huu-Tien, Dang, Nguyen, Minh-Phuong, Inoue, Naoya
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866917044351401984
author Huu-Tien, Dang
Nguyen, Minh-Phuong
Inoue, Naoya
author_facet Huu-Tien, Dang
Nguyen, Minh-Phuong
Inoue, Naoya
contents Label noise in datasets could significantly damage the performance and robustness of deep neural networks (DNNs) trained on these datasets. As the size of modern DNNs grows, there is a growing demand for automated tools for detecting such errors. In this paper, we propose post-hoc, model-agnostic noise detection and rectification methods utilizing the penultimate feature from a DNN. Our idea is based on the observation that the similarity between the penultimate feature of a mislabeled data point and its true class data points is higher than that for data points from other classes, making the probability of label occurrence within a tight, similar cluster informative for detecting and rectifying errors. Through theoretical and empirical analyses, we demonstrate that our approach achieves high detection performance across diverse, realistic noise scenarios and can automatically rectify these errors to improve dataset quality. Our implementation is available at https://anonymous.4open.science/r/noise-detection-and-rectification-AD8E.
format Preprint
id arxiv_https___arxiv_org_abs_2509_23964
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Detecting and Rectifying Noisy Labels: A Similarity-based Approach
Huu-Tien, Dang
Nguyen, Minh-Phuong
Inoue, Naoya
Machine Learning
Computation and Language
Label noise in datasets could significantly damage the performance and robustness of deep neural networks (DNNs) trained on these datasets. As the size of modern DNNs grows, there is a growing demand for automated tools for detecting such errors. In this paper, we propose post-hoc, model-agnostic noise detection and rectification methods utilizing the penultimate feature from a DNN. Our idea is based on the observation that the similarity between the penultimate feature of a mislabeled data point and its true class data points is higher than that for data points from other classes, making the probability of label occurrence within a tight, similar cluster informative for detecting and rectifying errors. Through theoretical and empirical analyses, we demonstrate that our approach achieves high detection performance across diverse, realistic noise scenarios and can automatically rectify these errors to improve dataset quality. Our implementation is available at https://anonymous.4open.science/r/noise-detection-and-rectification-AD8E.
title Detecting and Rectifying Noisy Labels: A Similarity-based Approach
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2509.23964