Data Preparation for Fairness-Performance Trade-Offs: A Practitioner-Friendly Alternative?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Voria, Gianmario, Di Matteo, Rebecca, Giordano, Giammaria, Catolino, Gemma, Palomba, Fabio
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913621604302848
author Voria, Gianmario
Di Matteo, Rebecca
Giordano, Giammaria
Catolino, Gemma
Palomba, Fabio
author_facet Voria, Gianmario
Di Matteo, Rebecca
Giordano, Giammaria
Catolino, Gemma
Palomba, Fabio
contents As machine learning (ML) systems are increasingly adopted across industries, addressing fairness and bias has become essential. While many solutions focus on ethical challenges in ML, recent studies highlight that data itself is a major source of bias. Pre-processing techniques, which mitigate bias before training, are effective but may impact model performance and pose integration difficulties. In contrast, fairness-aware Data Preparation practices are both familiar to practitioners and easier to implement, providing a more accessible approach to reducing bias. Objective. This registered report proposes an empirical evaluation of how optimally selected fairness-aware practices, applied in early ML lifecycle stages, can enhance both fairness and performance, potentially outperforming standard pre-processing bias mitigation methods. Method. To this end, we will introduce FATE, an optimization technique for selecting 'Data Preparation' pipelines that optimize fairness and performance. Using FATE, we will analyze the fairness-performance trade-off, comparing pipelines selected by FATE with results by pre-processing bias mitigation techniques.
format Preprint
id arxiv_https___arxiv_org_abs_2412_15920
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Data Preparation for Fairness-Performance Trade-Offs: A Practitioner-Friendly Alternative?
Voria, Gianmario
Di Matteo, Rebecca
Giordano, Giammaria
Catolino, Gemma
Palomba, Fabio
Software Engineering
Machine Learning
As machine learning (ML) systems are increasingly adopted across industries, addressing fairness and bias has become essential. While many solutions focus on ethical challenges in ML, recent studies highlight that data itself is a major source of bias. Pre-processing techniques, which mitigate bias before training, are effective but may impact model performance and pose integration difficulties. In contrast, fairness-aware Data Preparation practices are both familiar to practitioners and easier to implement, providing a more accessible approach to reducing bias. Objective. This registered report proposes an empirical evaluation of how optimally selected fairness-aware practices, applied in early ML lifecycle stages, can enhance both fairness and performance, potentially outperforming standard pre-processing bias mitigation methods. Method. To this end, we will introduce FATE, an optimization technique for selecting 'Data Preparation' pipelines that optimize fairness and performance. Using FATE, we will analyze the fairness-performance trade-off, comparing pipelines selected by FATE with results by pre-processing bias mitigation techniques.
title Data Preparation for Fairness-Performance Trade-Offs: A Practitioner-Friendly Alternative?
topic Software Engineering
Machine Learning
url https://arxiv.org/abs/2412.15920