Fairness Overfitting in Machine Learning: An Information-Theoretic Perspective

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Laakom, Firas, Chen, Haobo, Schmidhuber, Jürgen, Bu, Yuheng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916785715937280
author Laakom, Firas
Chen, Haobo
Schmidhuber, Jürgen
Bu, Yuheng
author_facet Laakom, Firas
Chen, Haobo
Schmidhuber, Jürgen
Bu, Yuheng
contents Despite substantial progress in promoting fairness in high-stake applications using machine learning models, existing methods often modify the training process, such as through regularizers or other interventions, but lack formal guarantees that fairness achieved during training will generalize to unseen data. Although overfitting with respect to prediction performance has been extensively studied, overfitting in terms of fairness loss has received far less attention. This paper proposes a theoretical framework for analyzing fairness generalization error through an information-theoretic lens. Our novel bounding technique is based on Efron-Stein inequality, which allows us to derive tight information-theoretic fairness generalization bounds with both Mutual Information (MI) and Conditional Mutual Information (CMI). Our empirical results validate the tightness and practical relevance of these bounds across diverse fairness-aware learning algorithms. Our framework offers valuable insights to guide the design of algorithms improving fairness generalization.
format Preprint
id arxiv_https___arxiv_org_abs_2506_07861
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Fairness Overfitting in Machine Learning: An Information-Theoretic Perspective
Laakom, Firas
Chen, Haobo
Schmidhuber, Jürgen
Bu, Yuheng
Machine Learning
Artificial Intelligence
Information Theory
Despite substantial progress in promoting fairness in high-stake applications using machine learning models, existing methods often modify the training process, such as through regularizers or other interventions, but lack formal guarantees that fairness achieved during training will generalize to unseen data. Although overfitting with respect to prediction performance has been extensively studied, overfitting in terms of fairness loss has received far less attention. This paper proposes a theoretical framework for analyzing fairness generalization error through an information-theoretic lens. Our novel bounding technique is based on Efron-Stein inequality, which allows us to derive tight information-theoretic fairness generalization bounds with both Mutual Information (MI) and Conditional Mutual Information (CMI). Our empirical results validate the tightness and practical relevance of these bounds across diverse fairness-aware learning algorithms. Our framework offers valuable insights to guide the design of algorithms improving fairness generalization.
title Fairness Overfitting in Machine Learning: An Information-Theoretic Perspective
topic Machine Learning
Artificial Intelligence
Information Theory
url https://arxiv.org/abs/2506.07861