Principled Evaluation with Human Labels: One Rater at a Time and Rater Equivalence

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Resnick, Paul, Kong, Yuqing, Schoenebeck, Grant, Weninger, Tim
Format: Preprint
Veröffentlicht: 2021
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914499651436544
author Resnick, Paul
Kong, Yuqing
Schoenebeck, Grant
Weninger, Tim
author_facet Resnick, Paul
Kong, Yuqing
Schoenebeck, Grant
Weninger, Tim
contents In many classification tasks, there is no definitive ground truth, only human judgments that may disagree. We address two challenges that arise in such settings: (1) how to use human raters to score classifiers, and (2) how to use them for comparison benchmarks. For the first, the common practice is to score classifiers against the majority vote of an evaluation panel of several human raters. We argue that this is not justified when either of two properties fails: objectivity or equanimity. Instead, under a utility model appropriate for such settings, scoring against one rater at a time and averaging the scores across raters is a more principled approach. For the second, we introduce the concept of rater equivalence: the smallest number of human raters whose combined judgment matches the classifier's performance. We provide a provably optimal algorithm for combining benchmark panel labels, and demonstrate the framework through case studies.
format Preprint
id arxiv_https___arxiv_org_abs_2106_01254
institution arXiv
publishDate 2021
record_format arxiv
spellingShingle Principled Evaluation with Human Labels: One Rater at a Time and Rater Equivalence
Resnick, Paul
Kong, Yuqing
Schoenebeck, Grant
Weninger, Tim
Machine Learning
Human-Computer Interaction
Multiagent Systems
In many classification tasks, there is no definitive ground truth, only human judgments that may disagree. We address two challenges that arise in such settings: (1) how to use human raters to score classifiers, and (2) how to use them for comparison benchmarks. For the first, the common practice is to score classifiers against the majority vote of an evaluation panel of several human raters. We argue that this is not justified when either of two properties fails: objectivity or equanimity. Instead, under a utility model appropriate for such settings, scoring against one rater at a time and averaging the scores across raters is a more principled approach. For the second, we introduce the concept of rater equivalence: the smallest number of human raters whose combined judgment matches the classifier's performance. We provide a provably optimal algorithm for combining benchmark panel labels, and demonstrate the framework through case studies.
title Principled Evaluation with Human Labels: One Rater at a Time and Rater Equivalence
topic Machine Learning
Human-Computer Interaction
Multiagent Systems
url https://arxiv.org/abs/2106.01254