Consistency is Key: Disentangling Label Variation in Natural Language Processing with Intra-Annotator Agreement

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Abercrombie, Gavin, Dinkar, Tanvi, Curry, Amanda Cercas, Rieser, Verena, Hovy, Dirk
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909854525816832
author Abercrombie, Gavin
Dinkar, Tanvi
Curry, Amanda Cercas
Rieser, Verena
Hovy, Dirk
author_facet Abercrombie, Gavin
Dinkar, Tanvi
Curry, Amanda Cercas
Rieser, Verena
Hovy, Dirk
contents We commonly use agreement measures to assess the utility of judgements made by human annotators in Natural Language Processing (NLP) tasks. While inter-annotator agreement is frequently used as an indication of label reliability by measuring consistency between annotators, we argue for the additional use of intra-annotator agreement to measure label stability (and annotator consistency) over time. However, in a systematic review, we find that the latter is rarely reported in this field. Calculating these measures can act as important quality control and could provide insights into why annotators disagree. We conduct exploratory annotation experiments to investigate the relationships between these measures and perceptions of subjectivity and ambiguity in text items, finding that annotators provide inconsistent responses around 25% of the time across four different NLP tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2301_10684
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Consistency is Key: Disentangling Label Variation in Natural Language Processing with Intra-Annotator Agreement
Abercrombie, Gavin
Dinkar, Tanvi
Curry, Amanda Cercas
Rieser, Verena
Hovy, Dirk
Computation and Language
We commonly use agreement measures to assess the utility of judgements made by human annotators in Natural Language Processing (NLP) tasks. While inter-annotator agreement is frequently used as an indication of label reliability by measuring consistency between annotators, we argue for the additional use of intra-annotator agreement to measure label stability (and annotator consistency) over time. However, in a systematic review, we find that the latter is rarely reported in this field. Calculating these measures can act as important quality control and could provide insights into why annotators disagree. We conduct exploratory annotation experiments to investigate the relationships between these measures and perceptions of subjectivity and ambiguity in text items, finding that annotators provide inconsistent responses around 25% of the time across four different NLP tasks.
title Consistency is Key: Disentangling Label Variation in Natural Language Processing with Intra-Annotator Agreement
topic Computation and Language
url https://arxiv.org/abs/2301.10684