Fair Dataset Distillation via Cross-Group Barycenter Alignment

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Moslemi, Mohammad Hossein, Dashtbayaz, Nima Hosseini, Mei, Zhimin, Ghaddar, Bissan, Wang, Boyu
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914584294588416
author Moslemi, Mohammad Hossein
Dashtbayaz, Nima Hosseini
Mei, Zhimin
Ghaddar, Bissan
Wang, Boyu
author_facet Moslemi, Mohammad Hossein
Dashtbayaz, Nima Hosseini
Mei, Zhimin
Ghaddar, Bissan
Wang, Boyu
contents Dataset Distillation aims to compress a large dataset into a small synthetic one while maintaining predictive performance. We show that as different demographic groups exhibit distinct predictive patterns, the distillation process struggles to simultaneously preserve informative signals for all subgroups, regardless of whether group sizes are mildly or severely imbalanced. Consequently, models trained on distilled data can experience substantial performance drops for certain subgroups, leading to fairness gaps. Crucially, these gaps do not disappear by merely correcting group imbalance, since they stem from fundamental mismatches in subgroup predictive patterns rather than from sample-size disparities alone. We therefore formally analyze the interaction between these two sources of bias and cast the solution as identifying a group-imbalance-agnostic barycenter of the predictive information that induces similar representations across all subgroups. By distilling toward this shared aggregate representation, we show that group fairness concerns can be reduced. Our approach is compatible with existing distillation methods, and empirical results show that it substantially reduces bias introduced by dataset distillation. Code is available at https://github.com/mhmoslemi/COBRA.
format Preprint
id arxiv_https___arxiv_org_abs_2605_00185
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Fair Dataset Distillation via Cross-Group Barycenter Alignment
Moslemi, Mohammad Hossein
Dashtbayaz, Nima Hosseini
Mei, Zhimin
Ghaddar, Bissan
Wang, Boyu
Machine Learning
Artificial Intelligence
Dataset Distillation aims to compress a large dataset into a small synthetic one while maintaining predictive performance. We show that as different demographic groups exhibit distinct predictive patterns, the distillation process struggles to simultaneously preserve informative signals for all subgroups, regardless of whether group sizes are mildly or severely imbalanced. Consequently, models trained on distilled data can experience substantial performance drops for certain subgroups, leading to fairness gaps. Crucially, these gaps do not disappear by merely correcting group imbalance, since they stem from fundamental mismatches in subgroup predictive patterns rather than from sample-size disparities alone. We therefore formally analyze the interaction between these two sources of bias and cast the solution as identifying a group-imbalance-agnostic barycenter of the predictive information that induces similar representations across all subgroups. By distilling toward this shared aggregate representation, we show that group fairness concerns can be reduced. Our approach is compatible with existing distillation methods, and empirical results show that it substantially reduces bias introduced by dataset distillation. Code is available at https://github.com/mhmoslemi/COBRA.
title Fair Dataset Distillation via Cross-Group Barycenter Alignment
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2605.00185