Data (in)equities in data science: Dissecting systemic and systematic biases in pulse oximetry

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Rountree, Lillian, Parikh, Harsh, Mukherjee, Bhramar
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913047793106944
author Rountree, Lillian
Parikh, Harsh
Mukherjee, Bhramar
author_facet Rountree, Lillian
Parikh, Harsh
Mukherjee, Bhramar
contents Data equity is an emerging framework for responsible data science. However, its core concepts, including fairness, representativeness, and information bias, remain largely abstract and general, lacking the mathematical specificity needed for practical implementation. In this paper, we demonstrate how statisticians can operationalize data equity by translating its tenets into precise, testable formulations tailored to a given problem. Using the well-documented case of differential measurement error across racial groups in pulse oximetry, we first adopt an oracle approach, tracing how a single upstream violation of information bias compounds through the analytic pipeline into treatment disparities, fairness violations, and adverse health outcomes. We then demonstrate the inverse: starting from an observed outcome disparity, the data equity framework provides a principled structure for systematically identifying its statistical sources. Our exposition reveals that data equity, prediction equity, and decision equity are distinct requirements with distinct evaluation and policy needs--a nuance that highlights both the unique role of statisticians in the era of artificial intelligence as well as the necessity of interdisciplinary collaboration.
format Preprint
id arxiv_https___arxiv_org_abs_2604_18291
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Data (in)equities in data science: Dissecting systemic and systematic biases in pulse oximetry
Rountree, Lillian
Parikh, Harsh
Mukherjee, Bhramar
Applications
Data equity is an emerging framework for responsible data science. However, its core concepts, including fairness, representativeness, and information bias, remain largely abstract and general, lacking the mathematical specificity needed for practical implementation. In this paper, we demonstrate how statisticians can operationalize data equity by translating its tenets into precise, testable formulations tailored to a given problem. Using the well-documented case of differential measurement error across racial groups in pulse oximetry, we first adopt an oracle approach, tracing how a single upstream violation of information bias compounds through the analytic pipeline into treatment disparities, fairness violations, and adverse health outcomes. We then demonstrate the inverse: starting from an observed outcome disparity, the data equity framework provides a principled structure for systematically identifying its statistical sources. Our exposition reveals that data equity, prediction equity, and decision equity are distinct requirements with distinct evaluation and policy needs--a nuance that highlights both the unique role of statisticians in the era of artificial intelligence as well as the necessity of interdisciplinary collaboration.
title Data (in)equities in data science: Dissecting systemic and systematic biases in pulse oximetry
topic Applications
url https://arxiv.org/abs/2604.18291