Data Preparation for Fairness-Performance Trade-Offs: A Practitioner-Friendly Alternative?
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866913621604302848 |
|---|---|
| author | Voria, Gianmario Di Matteo, Rebecca Giordano, Giammaria Catolino, Gemma Palomba, Fabio |
| author_facet | Voria, Gianmario Di Matteo, Rebecca Giordano, Giammaria Catolino, Gemma Palomba, Fabio |
| contents | As machine learning (ML) systems are increasingly adopted across industries, addressing fairness and bias has become essential. While many solutions focus on ethical challenges in ML, recent studies highlight that data itself is a major source of bias. Pre-processing techniques, which mitigate bias before training, are effective but may impact model performance and pose integration difficulties. In contrast, fairness-aware Data Preparation practices are both familiar to practitioners and easier to implement, providing a more accessible approach to reducing bias. Objective. This registered report proposes an empirical evaluation of how optimally selected fairness-aware practices, applied in early ML lifecycle stages, can enhance both fairness and performance, potentially outperforming standard pre-processing bias mitigation methods. Method. To this end, we will introduce FATE, an optimization technique for selecting 'Data Preparation' pipelines that optimize fairness and performance. Using FATE, we will analyze the fairness-performance trade-off, comparing pipelines selected by FATE with results by pre-processing bias mitigation techniques. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2412_15920 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Data Preparation for Fairness-Performance Trade-Offs: A Practitioner-Friendly Alternative? Voria, Gianmario Di Matteo, Rebecca Giordano, Giammaria Catolino, Gemma Palomba, Fabio Software Engineering Machine Learning As machine learning (ML) systems are increasingly adopted across industries, addressing fairness and bias has become essential. While many solutions focus on ethical challenges in ML, recent studies highlight that data itself is a major source of bias. Pre-processing techniques, which mitigate bias before training, are effective but may impact model performance and pose integration difficulties. In contrast, fairness-aware Data Preparation practices are both familiar to practitioners and easier to implement, providing a more accessible approach to reducing bias. Objective. This registered report proposes an empirical evaluation of how optimally selected fairness-aware practices, applied in early ML lifecycle stages, can enhance both fairness and performance, potentially outperforming standard pre-processing bias mitigation methods. Method. To this end, we will introduce FATE, an optimization technique for selecting 'Data Preparation' pipelines that optimize fairness and performance. Using FATE, we will analyze the fairness-performance trade-off, comparing pipelines selected by FATE with results by pre-processing bias mitigation techniques. |
| title | Data Preparation for Fairness-Performance Trade-Offs: A Practitioner-Friendly Alternative? |
| topic | Software Engineering Machine Learning |
| url | https://arxiv.org/abs/2412.15920 |