Managing Cognitive Bias in Human Labeling Operations for Rare-Event AI: Evidence from a Field Experiment

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Epping, Gunnar P., Caplin, Andrew, Duhaime, Erik, Holmes, William R., Martin, Daniel, Trueblood, Jennifer S.
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912963053486080
author Epping, Gunnar P.
Caplin, Andrew
Duhaime, Erik
Holmes, William R.
Martin, Daniel
Trueblood, Jennifer S.
author_facet Epping, Gunnar P.
Caplin, Andrew
Duhaime, Erik
Holmes, William R.
Martin, Daniel
Trueblood, Jennifer S.
contents Many operational AI systems depend on large-scale human annotation to detect rare but consequential events (e.g., fraud, defects, and medical abnormalities). When positives are rare, the prevalence effect induces systematic cognitive biases that inflate misses and can propagate through the AI lifecycle via biased training labels. We analyze prior experimental evidence and run a field experiment on DiagnosUs, a medical crowdsourcing platform, in which we hold the true prevalence in the unlabeled stream fixed (20% blasts) while varying (i) the prevalence of positives in the gold-standard feedback stream (20% vs. 50%) and (ii) the response interface (binary labels vs. elicited probabilities). We then post-process probabilistic labels using a linear-in-log-odds recalibration approach at the worker and crowd levels, and train convolutional neural networks on the resulting labels. Balanced feedback and probabilistic elicitation reduce rare-event misses, and pipeline-level recalibration substantially improves both classification performance and probabilistic calibration; these gains carry through to downstream CNN reliability out of sample.
format Preprint
id arxiv_https___arxiv_org_abs_2603_11511
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Managing Cognitive Bias in Human Labeling Operations for Rare-Event AI: Evidence from a Field Experiment
Epping, Gunnar P.
Caplin, Andrew
Duhaime, Erik
Holmes, William R.
Martin, Daniel
Trueblood, Jennifer S.
Human-Computer Interaction
General Economics
Economics
Many operational AI systems depend on large-scale human annotation to detect rare but consequential events (e.g., fraud, defects, and medical abnormalities). When positives are rare, the prevalence effect induces systematic cognitive biases that inflate misses and can propagate through the AI lifecycle via biased training labels. We analyze prior experimental evidence and run a field experiment on DiagnosUs, a medical crowdsourcing platform, in which we hold the true prevalence in the unlabeled stream fixed (20% blasts) while varying (i) the prevalence of positives in the gold-standard feedback stream (20% vs. 50%) and (ii) the response interface (binary labels vs. elicited probabilities). We then post-process probabilistic labels using a linear-in-log-odds recalibration approach at the worker and crowd levels, and train convolutional neural networks on the resulting labels. Balanced feedback and probabilistic elicitation reduce rare-event misses, and pipeline-level recalibration substantially improves both classification performance and probabilistic calibration; these gains carry through to downstream CNN reliability out of sample.
title Managing Cognitive Bias in Human Labeling Operations for Rare-Event AI: Evidence from a Field Experiment
topic Human-Computer Interaction
General Economics
Economics
url https://arxiv.org/abs/2603.11511