Optimal Fair Aggregation of Crowdsourced Noisy Labels using Demographic Parity Constraints

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Singer, Gabriel, Gruffaz, Samuel, Van, Olivier Vo, Vayatis, Nicolas, Kalogeratos, Argyris
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910006309289984
author Singer, Gabriel
Gruffaz, Samuel
Van, Olivier Vo
Vayatis, Nicolas
Kalogeratos, Argyris
author_facet Singer, Gabriel
Gruffaz, Samuel
Van, Olivier Vo
Vayatis, Nicolas
Kalogeratos, Argyris
contents As acquiring reliable ground-truth labels is usually costly, or infeasible, crowdsourcing and aggregation of noisy human annotations is the typical resort. Aggregating subjective labels, though, may amplify individual biases, particularly regarding sensitive features, raising fairness concerns. Nonetheless, fairness in crowdsourced aggregation remains largely unexplored, with no existing convergence guarantees and only limited post-processing approaches for enforcing $\varepsilon$-fairness under demographic parity. We address this gap by analyzing the fairness s of crowdsourced aggregation methods within the $\varepsilon$-fairness framework, for Majority Vote and Optimal Bayesian aggregation. In the small-crowd regime, we derive an upper bound on the fairness gap of Majority Vote in terms of the fairness gaps of the individual annotators. We further show that the fairness gap of the aggregated consensus converges exponentially fast to that of the ground-truth under interpretable conditions. Since ground-truth itself may still be unfair, we generalize a state-of-the-art multiclass fairness post-processing algorithm from the continuous to the discrete setting, which enforces strict demographic parity constraints to any aggregation rule. Experiments on synthetic and real datasets demonstrate the effectiveness of our approach and corroborate the theoretical insights.
format Preprint
id arxiv_https___arxiv_org_abs_2601_23221
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Optimal Fair Aggregation of Crowdsourced Noisy Labels using Demographic Parity Constraints
Singer, Gabriel
Gruffaz, Samuel
Van, Olivier Vo
Vayatis, Nicolas
Kalogeratos, Argyris
Machine Learning
As acquiring reliable ground-truth labels is usually costly, or infeasible, crowdsourcing and aggregation of noisy human annotations is the typical resort. Aggregating subjective labels, though, may amplify individual biases, particularly regarding sensitive features, raising fairness concerns. Nonetheless, fairness in crowdsourced aggregation remains largely unexplored, with no existing convergence guarantees and only limited post-processing approaches for enforcing $\varepsilon$-fairness under demographic parity. We address this gap by analyzing the fairness s of crowdsourced aggregation methods within the $\varepsilon$-fairness framework, for Majority Vote and Optimal Bayesian aggregation. In the small-crowd regime, we derive an upper bound on the fairness gap of Majority Vote in terms of the fairness gaps of the individual annotators. We further show that the fairness gap of the aggregated consensus converges exponentially fast to that of the ground-truth under interpretable conditions. Since ground-truth itself may still be unfair, we generalize a state-of-the-art multiclass fairness post-processing algorithm from the continuous to the discrete setting, which enforces strict demographic parity constraints to any aggregation rule. Experiments on synthetic and real datasets demonstrate the effectiveness of our approach and corroborate the theoretical insights.
title Optimal Fair Aggregation of Crowdsourced Noisy Labels using Demographic Parity Constraints
topic Machine Learning
url https://arxiv.org/abs/2601.23221