Exploiting Decision Trees to Detect Indirect Discrimination

Fuente: Zenodo
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Gouveia, Filipe, Lynce, Inês
Format: Recurso digital
Veröffentlicht: Zenodo 2022
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866901952940474368
author Gouveia, Filipe
Lynce, Inês
author_facet Gouveia, Filipe
Lynce, Inês
contents <p>Machine learning systems are increasingly present in our everyday lives, making decisions in multiple scenarios such as credit loans or advertisement displays. <br>For this reason, there is a crescent concern regarding fairness in applying machine learning systems. <br>Fairness is particularly relevant when such systems can affect a person's life.</p> <p>When training classification models, several approaches consider that some attributes are protected to avoid discriminatory behavior. Hence, a model should not use these attributes to make a decision.<br>These attributes are typically sensitive features for which it is desired not to have discriminatory behavior or bias (e.g., race or gender).<br>However, it may be possible to infer information of a protected attribute based on other non-protected attributes' data, leading to (indirect) discriminatory behavior even without using the protected attributes.</p> <p>In this work, we introduce an approach to assess whether a dataset used to train a machine learning model contains potential sources of indirect discrimination. This approach exploits the use of decision trees to identify such potential sources of discrimination.<br>We can identify from which attributes it is possible to infer information regarding a protected attribute. These attributes are called proxy attributes.<br>Moreover, we consider sets of proxy attributes from which it is possible to infer sensitive information when combined. If these attributes were deemed separately, no information could be inferred.</p> <p>The proposed approach is applied to several datasets used in fairness-awareness studies. The experimental evaluation identifies possible reasons for indirect discriminatory behavior in the datasets.<br>Moreover, the evaluation also considers thresholds that allow the identification of proxy attributes with a certain degree of certainty. Therefore, we can identify proxy attributes even when the dataset contains noisy data.</p>
format Recurso digital
id zenodo_https___doi_org_10_5281_zenodo_14810692
institution Zenodo
language
publishDate 2022
publisher Zenodo
record_format zenodo
spellingShingle Exploiting Decision Trees to Detect Indirect Discrimination
Gouveia, Filipe
Lynce, Inês
<p>Machine learning systems are increasingly present in our everyday lives, making decisions in multiple scenarios such as credit loans or advertisement displays. <br>For this reason, there is a crescent concern regarding fairness in applying machine learning systems. <br>Fairness is particularly relevant when such systems can affect a person's life.</p> <p>When training classification models, several approaches consider that some attributes are protected to avoid discriminatory behavior. Hence, a model should not use these attributes to make a decision.<br>These attributes are typically sensitive features for which it is desired not to have discriminatory behavior or bias (e.g., race or gender).<br>However, it may be possible to infer information of a protected attribute based on other non-protected attributes' data, leading to (indirect) discriminatory behavior even without using the protected attributes.</p> <p>In this work, we introduce an approach to assess whether a dataset used to train a machine learning model contains potential sources of indirect discrimination. This approach exploits the use of decision trees to identify such potential sources of discrimination.<br>We can identify from which attributes it is possible to infer information regarding a protected attribute. These attributes are called proxy attributes.<br>Moreover, we consider sets of proxy attributes from which it is possible to infer sensitive information when combined. If these attributes were deemed separately, no information could be inferred.</p> <p>The proposed approach is applied to several datasets used in fairness-awareness studies. The experimental evaluation identifies possible reasons for indirect discriminatory behavior in the datasets.<br>Moreover, the evaluation also considers thresholds that allow the identification of proxy attributes with a certain degree of certainty. Therefore, we can identify proxy attributes even when the dataset contains noisy data.</p>
title Exploiting Decision Trees to Detect Indirect Discrimination
url https://doi.org/10.5281/zenodo.14810692