Balancing Quality and Variation: Spam Filtering Distorts Data Label Distributions

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Fleisig, Eve, Orlikowski, Matthias, Cimiano, Philipp, Klein, Dan
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912691110543360
author Fleisig, Eve
Orlikowski, Matthias
Cimiano, Philipp
Klein, Dan
author_facet Fleisig, Eve
Orlikowski, Matthias
Cimiano, Philipp
Klein, Dan
contents For machine learning datasets to accurately represent diverse opinions in a population, they must preserve variation in data labels while filtering out spam or low-quality responses. How can we balance annotator reliability and representation? We empirically evaluate how a range of heuristics for annotator filtering affect the preservation of variation on subjective tasks. We find that these methods, designed for contexts in which variation from a single ground-truth label is considered noise, often remove annotators who disagree instead of spam annotators, introducing suboptimal tradeoffs between accuracy and label diversity. We find that conservative settings for annotator removal (<5%) are best, after which all tested methods increase the mean absolute error from the true average label. We analyze performance on synthetic spam to observe that these methods often assume spam annotators are more random than real spammers tend to be: most spammers are distributionally indistinguishable from real annotators, and the minority that are distinguishable tend to give relatively fixed answers, not random ones. Thus, tasks requiring the preservation of variation reverse the intuition of existing spam filtering methods: spammers tend to be less random than non-spammers, so metrics that assume variation is spam fare worse. These results highlight the need for spam removal methods that account for label diversity.
format Preprint
id arxiv_https___arxiv_org_abs_2509_08217
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Balancing Quality and Variation: Spam Filtering Distorts Data Label Distributions
Fleisig, Eve
Orlikowski, Matthias
Cimiano, Philipp
Klein, Dan
Computation and Language
Artificial Intelligence
For machine learning datasets to accurately represent diverse opinions in a population, they must preserve variation in data labels while filtering out spam or low-quality responses. How can we balance annotator reliability and representation? We empirically evaluate how a range of heuristics for annotator filtering affect the preservation of variation on subjective tasks. We find that these methods, designed for contexts in which variation from a single ground-truth label is considered noise, often remove annotators who disagree instead of spam annotators, introducing suboptimal tradeoffs between accuracy and label diversity. We find that conservative settings for annotator removal (<5%) are best, after which all tested methods increase the mean absolute error from the true average label. We analyze performance on synthetic spam to observe that these methods often assume spam annotators are more random than real spammers tend to be: most spammers are distributionally indistinguishable from real annotators, and the minority that are distinguishable tend to give relatively fixed answers, not random ones. Thus, tasks requiring the preservation of variation reverse the intuition of existing spam filtering methods: spammers tend to be less random than non-spammers, so metrics that assume variation is spam fare worse. These results highlight the need for spam removal methods that account for label diversity.
title Balancing Quality and Variation: Spam Filtering Distorts Data Label Distributions
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2509.08217