Bias Begins with Data: The FairGround Corpus for Robust and Reproducible Research on Algorithmic Fairness

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Simson, Jan, Fabris, Alessandro, Fröhner, Cosima, Kreuter, Frauke, Kern, Christoph
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918171314749440
author Simson, Jan
Fabris, Alessandro
Fröhner, Cosima
Kreuter, Frauke
Kern, Christoph
author_facet Simson, Jan
Fabris, Alessandro
Fröhner, Cosima
Kreuter, Frauke
Kern, Christoph
contents As machine learning (ML) systems are increasingly adopted in high-stakes decision-making domains, ensuring fairness in their outputs has become a central challenge. At the core of fair ML research are the datasets used to investigate bias and develop mitigation strategies. Yet, much of the existing work relies on a narrow selection of datasets--often arbitrarily chosen, inconsistently processed, and lacking in diversity--undermining the generalizability and reproducibility of results. To address these limitations, we present FairGround: a unified framework, data corpus, and Python package aimed at advancing reproducible research and critical data studies in fair ML classification. FairGround currently comprises 44 tabular datasets, each annotated with rich fairness-relevant metadata. Our accompanying Python package standardizes dataset loading, preprocessing, transformation, and splitting, streamlining experimental workflows. By providing a diverse and well-documented dataset corpus along with robust tooling, FairGround enables the development of fairer, more reliable, and more reproducible ML models. All resources are publicly available to support open and collaborative research.
format Preprint
id arxiv_https___arxiv_org_abs_2510_22363
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Bias Begins with Data: The FairGround Corpus for Robust and Reproducible Research on Algorithmic Fairness
Simson, Jan
Fabris, Alessandro
Fröhner, Cosima
Kreuter, Frauke
Kern, Christoph
Machine Learning
Computers and Society
As machine learning (ML) systems are increasingly adopted in high-stakes decision-making domains, ensuring fairness in their outputs has become a central challenge. At the core of fair ML research are the datasets used to investigate bias and develop mitigation strategies. Yet, much of the existing work relies on a narrow selection of datasets--often arbitrarily chosen, inconsistently processed, and lacking in diversity--undermining the generalizability and reproducibility of results. To address these limitations, we present FairGround: a unified framework, data corpus, and Python package aimed at advancing reproducible research and critical data studies in fair ML classification. FairGround currently comprises 44 tabular datasets, each annotated with rich fairness-relevant metadata. Our accompanying Python package standardizes dataset loading, preprocessing, transformation, and splitting, streamlining experimental workflows. By providing a diverse and well-documented dataset corpus along with robust tooling, FairGround enables the development of fairer, more reliable, and more reproducible ML models. All resources are publicly available to support open and collaborative research.
title Bias Begins with Data: The FairGround Corpus for Robust and Reproducible Research on Algorithmic Fairness
topic Machine Learning
Computers and Society
url https://arxiv.org/abs/2510.22363