Addressing Discretization-Induced Bias in Demographic Prediction

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Dong, Evan, Schein, Aaron, Wang, Yixin, Garg, Nikhil
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916261593612288
author Dong, Evan
Schein, Aaron
Wang, Yixin
Garg, Nikhil
author_facet Dong, Evan
Schein, Aaron
Wang, Yixin
Garg, Nikhil
contents Racial and other demographic imputation is necessary for many applications, especially in auditing disparities and outreach targeting in political campaigns. The canonical approach is to construct continuous predictions -- e.g., based on name and geography -- and then to $\textit{discretize}$ the predictions by selecting the most likely class (argmax). We study how this practice produces $\textit{discretization bias}$. In particular, we show that argmax labeling, as used by a prominent commercial voter file vendor to impute race/ethnicity, results in a substantial under-count of African-American voters, e.g., by 28.2% points in North Carolina. This bias can have substantial implications in downstream tasks that use such labels. We then introduce a $\textit{joint optimization}$ approach -- and a tractable $\textit{data-driven thresholding}$ heuristic -- that can eliminate this bias, with negligible individual-level accuracy loss. Finally, we theoretically analyze discretization bias, show that calibrated continuous models are insufficient to eliminate it, and that an approach such as ours is necessary. Broadly, we warn researchers and practitioners against discretizing continuous demographic predictions without considering downstream consequences.
format Preprint
id arxiv_https___arxiv_org_abs_2405_16762
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Addressing Discretization-Induced Bias in Demographic Prediction
Dong, Evan
Schein, Aaron
Wang, Yixin
Garg, Nikhil
Computers and Society
Machine Learning
K.4.0
Racial and other demographic imputation is necessary for many applications, especially in auditing disparities and outreach targeting in political campaigns. The canonical approach is to construct continuous predictions -- e.g., based on name and geography -- and then to $\textit{discretize}$ the predictions by selecting the most likely class (argmax). We study how this practice produces $\textit{discretization bias}$. In particular, we show that argmax labeling, as used by a prominent commercial voter file vendor to impute race/ethnicity, results in a substantial under-count of African-American voters, e.g., by 28.2% points in North Carolina. This bias can have substantial implications in downstream tasks that use such labels. We then introduce a $\textit{joint optimization}$ approach -- and a tractable $\textit{data-driven thresholding}$ heuristic -- that can eliminate this bias, with negligible individual-level accuracy loss. Finally, we theoretically analyze discretization bias, show that calibrated continuous models are insufficient to eliminate it, and that an approach such as ours is necessary. Broadly, we warn researchers and practitioners against discretizing continuous demographic predictions without considering downstream consequences.
title Addressing Discretization-Induced Bias in Demographic Prediction
topic Computers and Society
Machine Learning
K.4.0
url https://arxiv.org/abs/2405.16762