Corpus Considerations for Annotator Modeling and Scaling

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Sarumi, Olufunke O., Neuendorf, Béla, Plepi, Joan, Flek, Lucie, Schlötterer, Jörg, Welch, Charles
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911843184803840
author Sarumi, Olufunke O.
Neuendorf, Béla
Plepi, Joan
Flek, Lucie
Schlötterer, Jörg
Welch, Charles
author_facet Sarumi, Olufunke O.
Neuendorf, Béla
Plepi, Joan
Flek, Lucie
Schlötterer, Jörg
Welch, Charles
contents Recent trends in natural language processing research and annotation tasks affirm a paradigm shift from the traditional reliance on a single ground truth to a focus on individual perspectives, particularly in subjective tasks. In scenarios where annotation tasks are meant to encompass diversity, models that solely rely on the majority class labels may inadvertently disregard valuable minority perspectives. This oversight could result in the omission of crucial information and, in a broader context, risk disrupting the balance within larger ecosystems. As the landscape of annotator modeling unfolds with diverse representation techniques, it becomes imperative to investigate their effectiveness with the fine-grained features of the datasets in view. This study systematically explores various annotator modeling techniques and compares their performance across seven corpora. From our findings, we show that the commonly used user token model consistently outperforms more complex models. We introduce a composite embedding approach and show distinct differences in which model performs best as a function of the agreement with a given dataset. Our findings shed light on the relationship between corpus statistics and annotator modeling performance, which informs future work on corpus construction and perspectivist NLP.
format Preprint
id arxiv_https___arxiv_org_abs_2404_02340
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Corpus Considerations for Annotator Modeling and Scaling
Sarumi, Olufunke O.
Neuendorf, Béla
Plepi, Joan
Flek, Lucie
Schlötterer, Jörg
Welch, Charles
Computation and Language
F.2.2; I.2.7
Recent trends in natural language processing research and annotation tasks affirm a paradigm shift from the traditional reliance on a single ground truth to a focus on individual perspectives, particularly in subjective tasks. In scenarios where annotation tasks are meant to encompass diversity, models that solely rely on the majority class labels may inadvertently disregard valuable minority perspectives. This oversight could result in the omission of crucial information and, in a broader context, risk disrupting the balance within larger ecosystems. As the landscape of annotator modeling unfolds with diverse representation techniques, it becomes imperative to investigate their effectiveness with the fine-grained features of the datasets in view. This study systematically explores various annotator modeling techniques and compares their performance across seven corpora. From our findings, we show that the commonly used user token model consistently outperforms more complex models. We introduce a composite embedding approach and show distinct differences in which model performs best as a function of the agreement with a given dataset. Our findings shed light on the relationship between corpus statistics and annotator modeling performance, which informs future work on corpus construction and perspectivist NLP.
title Corpus Considerations for Annotator Modeling and Scaling
topic Computation and Language
F.2.2; I.2.7
url https://arxiv.org/abs/2404.02340