From Label Error Detection to Correction: A Modular Framework and Benchmark for Object Detection Datasets

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Penquitt, Sarina, Klees, Jonathan, Cakaj, Rinor, Kondermann, Daniel, Rottmann, Matthias, Schmarje, Lars
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911410046369792
author Penquitt, Sarina
Klees, Jonathan
Cakaj, Rinor
Kondermann, Daniel
Rottmann, Matthias
Schmarje, Lars
author_facet Penquitt, Sarina
Klees, Jonathan
Cakaj, Rinor
Kondermann, Daniel
Rottmann, Matthias
Schmarje, Lars
contents Object detection has advanced rapidly in recent years, driven by increasingly large and diverse datasets. However, label errors often compromise the quality of these datasets and affect the outcomes of training and benchmark evaluations. Although label error detection methods for object detection datasets now exist, they are typically validated only on synthetic benchmarks or via limited manual inspection. How to correct such errors systematically and at scale remains an open problem. We introduce a semi-automated framework for label error correction called Rechecked. Building on existing label error detection methods, their error proposals are reviewed with lightweight, crowd-sourced microtasks. We apply Rechecked to the class pedestrian in the KITTI dataset, for which we crowdsourced high-quality corrected annotations. We detect 18% of missing and inaccurate labels in the original ground truth. We show that current label error detection methods, when combined with our correction framework, can recover hundreds of errors with little human effort compared to annotation from scratch. However, even the best methods still miss up to 66% of the label errors, which motivates further research, now enabled by our released benchmark.
format Preprint
id arxiv_https___arxiv_org_abs_2508_06556
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle From Label Error Detection to Correction: A Modular Framework and Benchmark for Object Detection Datasets
Penquitt, Sarina
Klees, Jonathan
Cakaj, Rinor
Kondermann, Daniel
Rottmann, Matthias
Schmarje, Lars
Computer Vision and Pattern Recognition
Machine Learning
Object detection has advanced rapidly in recent years, driven by increasingly large and diverse datasets. However, label errors often compromise the quality of these datasets and affect the outcomes of training and benchmark evaluations. Although label error detection methods for object detection datasets now exist, they are typically validated only on synthetic benchmarks or via limited manual inspection. How to correct such errors systematically and at scale remains an open problem. We introduce a semi-automated framework for label error correction called Rechecked. Building on existing label error detection methods, their error proposals are reviewed with lightweight, crowd-sourced microtasks. We apply Rechecked to the class pedestrian in the KITTI dataset, for which we crowdsourced high-quality corrected annotations. We detect 18% of missing and inaccurate labels in the original ground truth. We show that current label error detection methods, when combined with our correction framework, can recover hundreds of errors with little human effort compared to annotation from scratch. However, even the best methods still miss up to 66% of the label errors, which motivates further research, now enabled by our released benchmark.
title From Label Error Detection to Correction: A Modular Framework and Benchmark for Object Detection Datasets
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2508.06556