Correcting Underrepresentation and Intersectional Bias for Classification

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Diana, Emily, Tolbert, Alexander Williams
Format: Preprint
Publié: 2023
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866910469975965696
author Diana, Emily
Tolbert, Alexander Williams
author_facet Diana, Emily
Tolbert, Alexander Williams
contents We consider the problem of learning from data corrupted by underrepresentation bias, where positive examples are filtered from the data at different, unknown rates for a fixed number of sensitive groups. We show that with a small amount of unbiased data, we can efficiently estimate the group-wise drop-out rates, even in settings where intersectional group membership makes learning each intersectional rate computationally infeasible. Using these estimates, we construct a reweighting scheme that allows us to approximate the loss of any hypothesis on the true distribution, even if we only observe the empirical error on a biased sample. From this, we present an algorithm encapsulating this learning and reweighting process along with a thorough empirical investigation. Finally, we define a bespoke notion of PAC learnability for the underrepresentation and intersectional bias setting and show that our algorithm permits efficient learning for model classes of finite VC dimension.
format Preprint
id arxiv_https___arxiv_org_abs_2306_11112
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Correcting Underrepresentation and Intersectional Bias for Classification
Diana, Emily
Tolbert, Alexander Williams
Machine Learning
Computers and Society
Data Structures and Algorithms
We consider the problem of learning from data corrupted by underrepresentation bias, where positive examples are filtered from the data at different, unknown rates for a fixed number of sensitive groups. We show that with a small amount of unbiased data, we can efficiently estimate the group-wise drop-out rates, even in settings where intersectional group membership makes learning each intersectional rate computationally infeasible. Using these estimates, we construct a reweighting scheme that allows us to approximate the loss of any hypothesis on the true distribution, even if we only observe the empirical error on a biased sample. From this, we present an algorithm encapsulating this learning and reweighting process along with a thorough empirical investigation. Finally, we define a bespoke notion of PAC learnability for the underrepresentation and intersectional bias setting and show that our algorithm permits efficient learning for model classes of finite VC dimension.
title Correcting Underrepresentation and Intersectional Bias for Classification
topic Machine Learning
Computers and Society
Data Structures and Algorithms
url https://arxiv.org/abs/2306.11112