Machine Learning with Multitype Protected Attributes: Intersectional Fairness through Regularisation

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Lee, Ho Ming, Antonio, Katrien, Avanzi, Benjamin, Marchi, Lorenzo, Zhou, Rui
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866917006476836864
author Lee, Ho Ming
Antonio, Katrien
Avanzi, Benjamin
Marchi, Lorenzo
Zhou, Rui
author_facet Lee, Ho Ming
Antonio, Katrien
Avanzi, Benjamin
Marchi, Lorenzo
Zhou, Rui
contents Ensuring equitable treatment (fairness) across protected attributes (such as gender or ethnicity) is a critical issue in machine learning. Most existing literature focuses on binary classification, but achieving fairness in regression tasks-such as insurance pricing or hiring score assessments-is equally important. Moreover, anti-discrimination laws also apply to continuous attributes, such as age, for which many existing methods are not applicable. In practice, multiple protected attributes can exist simultaneously; however, methods targeting fairness across several attributes often overlook so-called "fairness gerrymandering", thereby ignoring disparities among intersectional subgroups (e.g., African-American women or Hispanic men). In this paper, we propose a distance covariance regularisation framework that mitigates the association between model predictions and protected attributes, in line with the fairness definition of demographic parity, and that captures both linear and nonlinear dependencies. To enhance applicability in the presence of multiple protected attributes, we extend our framework by incorporating two multivariate dependence measures based on distance covariance: the previously proposed joint distance covariance (JdCov) and our novel concatenated distance covariance (CCdCov), which effectively address fairness gerrymandering in both regression and classification tasks involving protected attributes of various types. We discuss and illustrate how to calibrate regularisation strength, including a method based on Jensen-Shannon divergence, which quantifies dissimilarities in prediction distributions across groups. We apply our framework to the COMPAS recidivism dataset and a large motor insurance claims dataset.
format Preprint
id arxiv_https___arxiv_org_abs_2509_08163
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Machine Learning with Multitype Protected Attributes: Intersectional Fairness through Regularisation
Lee, Ho Ming
Antonio, Katrien
Avanzi, Benjamin
Marchi, Lorenzo
Zhou, Rui
Machine Learning
Risk Management
Applications
90B50, 62P05, 62H20, 68T07
Ensuring equitable treatment (fairness) across protected attributes (such as gender or ethnicity) is a critical issue in machine learning. Most existing literature focuses on binary classification, but achieving fairness in regression tasks-such as insurance pricing or hiring score assessments-is equally important. Moreover, anti-discrimination laws also apply to continuous attributes, such as age, for which many existing methods are not applicable. In practice, multiple protected attributes can exist simultaneously; however, methods targeting fairness across several attributes often overlook so-called "fairness gerrymandering", thereby ignoring disparities among intersectional subgroups (e.g., African-American women or Hispanic men). In this paper, we propose a distance covariance regularisation framework that mitigates the association between model predictions and protected attributes, in line with the fairness definition of demographic parity, and that captures both linear and nonlinear dependencies. To enhance applicability in the presence of multiple protected attributes, we extend our framework by incorporating two multivariate dependence measures based on distance covariance: the previously proposed joint distance covariance (JdCov) and our novel concatenated distance covariance (CCdCov), which effectively address fairness gerrymandering in both regression and classification tasks involving protected attributes of various types. We discuss and illustrate how to calibrate regularisation strength, including a method based on Jensen-Shannon divergence, which quantifies dissimilarities in prediction distributions across groups. We apply our framework to the COMPAS recidivism dataset and a large motor insurance claims dataset.
title Machine Learning with Multitype Protected Attributes: Intersectional Fairness through Regularisation
topic Machine Learning
Risk Management
Applications
90B50, 62P05, 62H20, 68T07
url https://arxiv.org/abs/2509.08163