Are we chasing ghosts? Quantifying unattributable polarization, and attributing the rest to annotator groups

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Tsirmpas, Dimitris, Pavlopoulos, John
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866917545939828736
author Tsirmpas, Dimitris
Pavlopoulos, John
author_facet Tsirmpas, Dimitris
Pavlopoulos, John
contents Standard agreement metrics often fail to capture systematic differences in opinion between minority and majority-group annotators, jeopardizing tasks such as hate speech and toxicity detection. Polarization has recently been proposed as a more robust way of distinguishing minor disagreements from systematic differences in opinion, but existing approaches do not provide practical tools for attributing it to specific annotator groups. We evaluate current methods and identify two major limitations in realistic settings: (1) the presence of ``inherent'' polarization that cannot be attributed to any known or latent groups, and (2) opposing polarization effects canceling each other out in aggregated annotations. To address these issues, we introduce a new metric that measures and tests the statistical significance of polarization attribution for annotator groups while avoiding these limitations, as well as an open-source Python library implementation, finding that no more than 20 annotators are needed per comment for reliable estimation. We apply our method to four subjective NLP datasets and find that gender and race consistently explain polarization patterns, while differences between annotator groups become stronger as the groups are further apart.
format Preprint
id arxiv_https___arxiv_org_abs_2602_06055
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Are we chasing ghosts? Quantifying unattributable polarization, and attributing the rest to annotator groups
Tsirmpas, Dimitris
Pavlopoulos, John
Computation and Language
68T09 (Primary)
G.3
Standard agreement metrics often fail to capture systematic differences in opinion between minority and majority-group annotators, jeopardizing tasks such as hate speech and toxicity detection. Polarization has recently been proposed as a more robust way of distinguishing minor disagreements from systematic differences in opinion, but existing approaches do not provide practical tools for attributing it to specific annotator groups. We evaluate current methods and identify two major limitations in realistic settings: (1) the presence of ``inherent'' polarization that cannot be attributed to any known or latent groups, and (2) opposing polarization effects canceling each other out in aggregated annotations. To address these issues, we introduce a new metric that measures and tests the statistical significance of polarization attribution for annotator groups while avoiding these limitations, as well as an open-source Python library implementation, finding that no more than 20 annotators are needed per comment for reliable estimation. We apply our method to four subjective NLP datasets and find that gender and race consistently explain polarization patterns, while differences between annotator groups become stronger as the groups are further apart.
title Are we chasing ghosts? Quantifying unattributable polarization, and attributing the rest to annotator groups
topic Computation and Language
68T09 (Primary)
G.3
url https://arxiv.org/abs/2602.06055