When majority rules, minority loses: bias amplification of gradient descent

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bachoc, François, Bolte, Jérôme, Boustany, Ryan, Loubes, Jean-Michel
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912658691719168
author Bachoc, François
Bolte, Jérôme
Boustany, Ryan
Loubes, Jean-Michel
author_facet Bachoc, François
Bolte, Jérôme
Boustany, Ryan
Loubes, Jean-Michel
contents Despite growing empirical evidence of bias amplification in machine learning, its theoretical foundations remain poorly understood. We develop a formal framework for majority-minority learning tasks, showing how standard training can favor majority groups and produce stereotypical predictors that neglect minority-specific features. Assuming population and variance imbalance, our analysis reveals three key findings: (i) the close proximity between ``full-data'' and stereotypical predictors, (ii) the dominance of a region where training the entire model tends to merely learn the majority traits, and (iii) a lower bound on the additional training required. Our results are illustrated through experiments in deep learning for tabular and image classification tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2505_13122
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle When majority rules, minority loses: bias amplification of gradient descent
Bachoc, François
Bolte, Jérôme
Boustany, Ryan
Loubes, Jean-Michel
Machine Learning
Artificial Intelligence
Optimization and Control
Despite growing empirical evidence of bias amplification in machine learning, its theoretical foundations remain poorly understood. We develop a formal framework for majority-minority learning tasks, showing how standard training can favor majority groups and produce stereotypical predictors that neglect minority-specific features. Assuming population and variance imbalance, our analysis reveals three key findings: (i) the close proximity between ``full-data'' and stereotypical predictors, (ii) the dominance of a region where training the entire model tends to merely learn the majority traits, and (iii) a lower bound on the additional training required. Our results are illustrated through experiments in deep learning for tabular and image classification tasks.
title When majority rules, minority loses: bias amplification of gradient descent
topic Machine Learning
Artificial Intelligence
Optimization and Control
url https://arxiv.org/abs/2505.13122