Learning to Detect Label Errors by Making Them: A Method for Segmentation and Object Detection Datasets

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Penquitt, Sarina, Riedlinger, Tobias, Heller, Timo, Reischl, Markus, Rottmann, Matthias
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908501940371456
author Penquitt, Sarina
Riedlinger, Tobias
Heller, Timo
Reischl, Markus
Rottmann, Matthias
author_facet Penquitt, Sarina
Riedlinger, Tobias
Heller, Timo
Reischl, Markus
Rottmann, Matthias
contents Recently, detection of label errors and improvement of label quality in datasets for supervised learning tasks has become an increasingly important goal in both research and industry. The consequences of incorrectly annotated data include reduced model performance, biased benchmark results, and lower overall accuracy. Current state-of-the-art label error detection methods often focus on a single computer vision task and, consequently, a specific type of dataset, containing, for example, either bounding boxes or pixel-wise annotations. Furthermore, previous methods are not learning-based. In this work, we overcome this research gap. We present a unified method for detecting label errors in object detection, semantic segmentation, and instance segmentation datasets. In a nutshell, our approach - learning to detect label errors by making them - works as follows: we inject different kinds of label errors into the ground truth. Then, the detection of label errors, across all mentioned primary tasks, is framed as an instance segmentation problem based on a composite input. In our experiments, we compare the label error detection performance of our method with various baselines and state-of-the-art approaches of each task's domain on simulated label errors across multiple tasks, datasets, and base models. This is complemented by a generalization study on real-world label errors. Additionally, we release 459 real label errors identified in the Cityscapes dataset and provide a benchmark for real label error detection in Cityscapes.
format Preprint
id arxiv_https___arxiv_org_abs_2508_17930
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Learning to Detect Label Errors by Making Them: A Method for Segmentation and Object Detection Datasets
Penquitt, Sarina
Riedlinger, Tobias
Heller, Timo
Reischl, Markus
Rottmann, Matthias
Machine Learning
Computer Vision and Pattern Recognition
Recently, detection of label errors and improvement of label quality in datasets for supervised learning tasks has become an increasingly important goal in both research and industry. The consequences of incorrectly annotated data include reduced model performance, biased benchmark results, and lower overall accuracy. Current state-of-the-art label error detection methods often focus on a single computer vision task and, consequently, a specific type of dataset, containing, for example, either bounding boxes or pixel-wise annotations. Furthermore, previous methods are not learning-based. In this work, we overcome this research gap. We present a unified method for detecting label errors in object detection, semantic segmentation, and instance segmentation datasets. In a nutshell, our approach - learning to detect label errors by making them - works as follows: we inject different kinds of label errors into the ground truth. Then, the detection of label errors, across all mentioned primary tasks, is framed as an instance segmentation problem based on a composite input. In our experiments, we compare the label error detection performance of our method with various baselines and state-of-the-art approaches of each task's domain on simulated label errors across multiple tasks, datasets, and base models. This is complemented by a generalization study on real-world label errors. Additionally, we release 459 real label errors identified in the Cityscapes dataset and provide a benchmark for real label error detection in Cityscapes.
title Learning to Detect Label Errors by Making Them: A Method for Segmentation and Object Detection Datasets
topic Machine Learning
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2508.17930