Efficient Online Crowdsourcing with Complex Annotations

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Meir, Reshef, Nguyen, Viet-An, Chen, Xu, Ramakrishnan, Jagdish, Weinsberg, Udi
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917576856043520
author Meir, Reshef
Nguyen, Viet-An
Chen, Xu
Ramakrishnan, Jagdish
Weinsberg, Udi
author_facet Meir, Reshef
Nguyen, Viet-An
Chen, Xu
Ramakrishnan, Jagdish
Weinsberg, Udi
contents Crowdsourcing platforms use various truth discovery algorithms to aggregate annotations from multiple labelers. In an online setting, however, the main challenge is to decide whether to ask for more annotations for each item to efficiently trade off cost (i.e., the number of annotations) for quality of the aggregated annotations. In this paper, we propose a novel approach for general complex annotation (such as bounding boxes and taxonomy paths), that works in an online crowdsourcing setting. We prove that the expected average similarity of a labeler is linear in their accuracy \emph{conditional on the reported label}. This enables us to infer reported label accuracy in a broad range of scenarios. We conduct extensive evaluations on real-world crowdsourcing data from Meta and show the effectiveness of our proposed online algorithms in improving the cost-quality trade-off.
format Preprint
id arxiv_https___arxiv_org_abs_2401_15116
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Efficient Online Crowdsourcing with Complex Annotations
Meir, Reshef
Nguyen, Viet-An
Chen, Xu
Ramakrishnan, Jagdish
Weinsberg, Udi
Human-Computer Interaction
Machine Learning
Crowdsourcing platforms use various truth discovery algorithms to aggregate annotations from multiple labelers. In an online setting, however, the main challenge is to decide whether to ask for more annotations for each item to efficiently trade off cost (i.e., the number of annotations) for quality of the aggregated annotations. In this paper, we propose a novel approach for general complex annotation (such as bounding boxes and taxonomy paths), that works in an online crowdsourcing setting. We prove that the expected average similarity of a labeler is linear in their accuracy \emph{conditional on the reported label}. This enables us to infer reported label accuracy in a broad range of scenarios. We conduct extensive evaluations on real-world crowdsourcing data from Meta and show the effectiveness of our proposed online algorithms in improving the cost-quality trade-off.
title Efficient Online Crowdsourcing with Complex Annotations
topic Human-Computer Interaction
Machine Learning
url https://arxiv.org/abs/2401.15116