Class-wise Autoencoders Measure Classification Difficulty And Detect Label Mistakes

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Marks, Jacob, Griffin, Brent A., Corso, Jason J.
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929613448413184
author Marks, Jacob
Griffin, Brent A.
Corso, Jason J.
author_facet Marks, Jacob
Griffin, Brent A.
Corso, Jason J.
contents We introduce a new framework for analyzing classification datasets based on the ratios of reconstruction errors between autoencoders trained on individual classes. This analysis framework enables efficient characterization of datasets on the sample, class, and entire dataset levels. We define reconstruction error ratios (RERs) that probe classification difficulty and allow its decomposition into (1) finite sample size and (2) Bayes error and decision-boundary complexity. Through systematic study across 19 popular visual datasets, we find that our RER-based dataset difficulty probe strongly correlates with error rate for state-of-the-art (SOTA) classification models. By interpreting sample-level classification difficulty as a label mistakenness score, we further find that RERs achieve SOTA performance on mislabel detection tasks on hard datasets under symmetric and asymmetric label noise. Our code is publicly available at https://github.com/voxel51/reconstruction-error-ratios.
format Preprint
id arxiv_https___arxiv_org_abs_2412_02596
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Class-wise Autoencoders Measure Classification Difficulty And Detect Label Mistakes
Marks, Jacob
Griffin, Brent A.
Corso, Jason J.
Machine Learning
Computer Vision and Pattern Recognition
We introduce a new framework for analyzing classification datasets based on the ratios of reconstruction errors between autoencoders trained on individual classes. This analysis framework enables efficient characterization of datasets on the sample, class, and entire dataset levels. We define reconstruction error ratios (RERs) that probe classification difficulty and allow its decomposition into (1) finite sample size and (2) Bayes error and decision-boundary complexity. Through systematic study across 19 popular visual datasets, we find that our RER-based dataset difficulty probe strongly correlates with error rate for state-of-the-art (SOTA) classification models. By interpreting sample-level classification difficulty as a label mistakenness score, we further find that RERs achieve SOTA performance on mislabel detection tasks on hard datasets under symmetric and asymmetric label noise. Our code is publicly available at https://github.com/voxel51/reconstruction-error-ratios.
title Class-wise Autoencoders Measure Classification Difficulty And Detect Label Mistakes
topic Machine Learning
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.02596