Robust Assignment of Labels for Active Learning with Sparse and Noisy Annotations

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Kałuża, Daniel, Janusz, Andrzej, Ślęzak, Dominik
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911944803352576
author Kałuża, Daniel
Janusz, Andrzej
Ślęzak, Dominik
author_facet Kałuża, Daniel
Janusz, Andrzej
Ślęzak, Dominik
contents Supervised classification algorithms are used to solve a growing number of real-life problems around the globe. Their performance is strictly connected with the quality of labels used in training. Unfortunately, acquiring good-quality annotations for many tasks is infeasible or too expensive to be done in practice. To tackle this challenge, active learning algorithms are commonly employed to select only the most relevant data for labeling. However, this is possible only when the quality and quantity of labels acquired from experts are sufficient. Unfortunately, in many applications, a trade-off between annotating individual samples by multiple annotators to increase label quality vs. annotating new samples to increase the total number of labeled instances is necessary. In this paper, we address the issue of faulty data annotations in the context of active learning. In particular, we propose two novel annotation unification algorithms that utilize unlabeled parts of the sample space. The proposed methods require little to no intersection between samples annotated by different experts. Our experiments on four public datasets indicate the robustness and superiority of the proposed methods in both, the estimation of the annotator's reliability, and the assignment of actual labels, against the state-of-the-art algorithms and the simple majority voting.
format Preprint
id arxiv_https___arxiv_org_abs_2307_14380
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Robust Assignment of Labels for Active Learning with Sparse and Noisy Annotations
Kałuża, Daniel
Janusz, Andrzej
Ślęzak, Dominik
Machine Learning
Human-Computer Interaction
I.2.6
Supervised classification algorithms are used to solve a growing number of real-life problems around the globe. Their performance is strictly connected with the quality of labels used in training. Unfortunately, acquiring good-quality annotations for many tasks is infeasible or too expensive to be done in practice. To tackle this challenge, active learning algorithms are commonly employed to select only the most relevant data for labeling. However, this is possible only when the quality and quantity of labels acquired from experts are sufficient. Unfortunately, in many applications, a trade-off between annotating individual samples by multiple annotators to increase label quality vs. annotating new samples to increase the total number of labeled instances is necessary. In this paper, we address the issue of faulty data annotations in the context of active learning. In particular, we propose two novel annotation unification algorithms that utilize unlabeled parts of the sample space. The proposed methods require little to no intersection between samples annotated by different experts. Our experiments on four public datasets indicate the robustness and superiority of the proposed methods in both, the estimation of the annotator's reliability, and the assignment of actual labels, against the state-of-the-art algorithms and the simple majority voting.
title Robust Assignment of Labels for Active Learning with Sparse and Noisy Annotations
topic Machine Learning
Human-Computer Interaction
I.2.6
url https://arxiv.org/abs/2307.14380