Metrics for Dataset Demographic Bias: A Case Study on Facial Expression Recognition

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Dominguez-Catena, Iris, Paternain, Daniel, Galar, Mikel
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929399193927680
author Dominguez-Catena, Iris
Paternain, Daniel
Galar, Mikel
author_facet Dominguez-Catena, Iris
Paternain, Daniel
Galar, Mikel
contents Demographic biases in source datasets have been shown as one of the causes of unfairness and discrimination in the predictions of Machine Learning models. One of the most prominent types of demographic bias are statistical imbalances in the representation of demographic groups in the datasets. In this paper, we study the measurement of these biases by reviewing the existing metrics, including those that can be borrowed from other disciplines. We develop a taxonomy for the classification of these metrics, providing a practical guide for the selection of appropriate metrics. To illustrate the utility of our framework, and to further understand the practical characteristics of the metrics, we conduct a case study of 20 datasets used in Facial Emotion Recognition (FER), analyzing the biases present in them. Our experimental results show that many metrics are redundant and that a reduced subset of metrics may be sufficient to measure the amount of demographic bias. The paper provides valuable insights for researchers in AI and related fields to mitigate dataset bias and improve the fairness and accuracy of AI models. The code is available at https://github.com/irisdominguez/dataset_bias_metrics.
format Preprint
id arxiv_https___arxiv_org_abs_2303_15889
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Metrics for Dataset Demographic Bias: A Case Study on Facial Expression Recognition
Dominguez-Catena, Iris
Paternain, Daniel
Galar, Mikel
Computer Vision and Pattern Recognition
Computers and Society
Demographic biases in source datasets have been shown as one of the causes of unfairness and discrimination in the predictions of Machine Learning models. One of the most prominent types of demographic bias are statistical imbalances in the representation of demographic groups in the datasets. In this paper, we study the measurement of these biases by reviewing the existing metrics, including those that can be borrowed from other disciplines. We develop a taxonomy for the classification of these metrics, providing a practical guide for the selection of appropriate metrics. To illustrate the utility of our framework, and to further understand the practical characteristics of the metrics, we conduct a case study of 20 datasets used in Facial Emotion Recognition (FER), analyzing the biases present in them. Our experimental results show that many metrics are redundant and that a reduced subset of metrics may be sufficient to measure the amount of demographic bias. The paper provides valuable insights for researchers in AI and related fields to mitigate dataset bias and improve the fairness and accuracy of AI models. The code is available at https://github.com/irisdominguez/dataset_bias_metrics.
title Metrics for Dataset Demographic Bias: A Case Study on Facial Expression Recognition
topic Computer Vision and Pattern Recognition
Computers and Society
url https://arxiv.org/abs/2303.15889