On Efficient and Statistical Quality Estimation for Data Annotation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Klie, Jan-Christoph, Haladjian, Juan, Kirchner, Marc, Nair, Rahul
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917677708083200
author Klie, Jan-Christoph
Haladjian, Juan
Kirchner, Marc
Nair, Rahul
author_facet Klie, Jan-Christoph
Haladjian, Juan
Kirchner, Marc
Nair, Rahul
contents Annotated datasets are an essential ingredient to train, evaluate, compare and productionalize supervised machine learning models. It is therefore imperative that annotations are of high quality. For their creation, good quality management and thereby reliable quality estimates are needed. Then, if quality is insufficient during the annotation process, rectifying measures can be taken to improve it. Quality estimation is often performed by having experts manually label instances as correct or incorrect. But checking all annotated instances tends to be expensive. Therefore, in practice, usually only subsets are inspected; sizes are chosen mostly without justification or regard to statistical power and more often than not, are relatively small. Basing estimates on small sample sizes, however, can lead to imprecise values for the error rate. Using unnecessarily large sample sizes costs money that could be better spent, for instance on more annotations. Therefore, we first describe in detail how to use confidence intervals for finding the minimal sample size needed to estimate the annotation error rate. Then, we propose applying acceptance sampling as an alternative to error rate estimation We show that acceptance sampling can reduce the required sample sizes up to 50% while providing the same statistical guarantees.
format Preprint
id arxiv_https___arxiv_org_abs_2405_11919
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle On Efficient and Statistical Quality Estimation for Data Annotation
Klie, Jan-Christoph
Haladjian, Juan
Kirchner, Marc
Nair, Rahul
Machine Learning
Artificial Intelligence
Computation and Language
Annotated datasets are an essential ingredient to train, evaluate, compare and productionalize supervised machine learning models. It is therefore imperative that annotations are of high quality. For their creation, good quality management and thereby reliable quality estimates are needed. Then, if quality is insufficient during the annotation process, rectifying measures can be taken to improve it. Quality estimation is often performed by having experts manually label instances as correct or incorrect. But checking all annotated instances tends to be expensive. Therefore, in practice, usually only subsets are inspected; sizes are chosen mostly without justification or regard to statistical power and more often than not, are relatively small. Basing estimates on small sample sizes, however, can lead to imprecise values for the error rate. Using unnecessarily large sample sizes costs money that could be better spent, for instance on more annotations. Therefore, we first describe in detail how to use confidence intervals for finding the minimal sample size needed to estimate the annotation error rate. Then, we propose applying acceptance sampling as an alternative to error rate estimation We show that acceptance sampling can reduce the required sample sizes up to 50% while providing the same statistical guarantees.
title On Efficient and Statistical Quality Estimation for Data Annotation
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2405.11919