Investigation of Accuracy and Bias in Face Recognition Trained with Synthetic Data

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Korshunov, Pavel, Kotwal, Ketan, Ecabert, Christophe, Vidit, Vidit, Mohammadi, Amir, Marcel, Sebastien
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913962397794304
author Korshunov, Pavel
Kotwal, Ketan
Ecabert, Christophe
Vidit, Vidit
Mohammadi, Amir
Marcel, Sebastien
author_facet Korshunov, Pavel
Kotwal, Ketan
Ecabert, Christophe
Vidit, Vidit
Mohammadi, Amir
Marcel, Sebastien
contents Synthetic data has emerged as a promising alternative for training face recognition (FR) models, offering advantages in scalability, privacy compliance, and potential for bias mitigation. However, critical questions remain on whether both high accuracy and fairness can be achieved with synthetic data. In this work, we evaluate the impact of synthetic data on bias and performance of FR systems. We generate balanced face dataset, FairFaceGen, using two state of the art text-to-image generators, Flux.1-dev and Stable Diffusion v3.5 (SD35), and combine them with several identity augmentation methods, including Arc2Face and four IP-Adapters. By maintaining equal identity count across synthetic and real datasets, we ensure fair comparisons when evaluating FR performance on standard (LFW, AgeDB-30, etc.) and challenging IJB-B/C benchmarks and FR bias on Racial Faces in-the-Wild (RFW) dataset. Our results demonstrate that although synthetic data still lags behind the real datasets in the generalization on IJB-B/C, demographically balanced synthetic datasets, especially those generated with SD35, show potential for bias mitigation. We also observe that the number and quality of intra-class augmentations significantly affect FR accuracy and fairness. These findings provide practical guidelines for constructing fairer FR systems using synthetic data.
format Preprint
id arxiv_https___arxiv_org_abs_2507_20782
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Investigation of Accuracy and Bias in Face Recognition Trained with Synthetic Data
Korshunov, Pavel
Kotwal, Ketan
Ecabert, Christophe
Vidit, Vidit
Mohammadi, Amir
Marcel, Sebastien
Computer Vision and Pattern Recognition
Artificial Intelligence
Synthetic data has emerged as a promising alternative for training face recognition (FR) models, offering advantages in scalability, privacy compliance, and potential for bias mitigation. However, critical questions remain on whether both high accuracy and fairness can be achieved with synthetic data. In this work, we evaluate the impact of synthetic data on bias and performance of FR systems. We generate balanced face dataset, FairFaceGen, using two state of the art text-to-image generators, Flux.1-dev and Stable Diffusion v3.5 (SD35), and combine them with several identity augmentation methods, including Arc2Face and four IP-Adapters. By maintaining equal identity count across synthetic and real datasets, we ensure fair comparisons when evaluating FR performance on standard (LFW, AgeDB-30, etc.) and challenging IJB-B/C benchmarks and FR bias on Racial Faces in-the-Wild (RFW) dataset. Our results demonstrate that although synthetic data still lags behind the real datasets in the generalization on IJB-B/C, demographically balanced synthetic datasets, especially those generated with SD35, show potential for bias mitigation. We also observe that the number and quality of intra-class augmentations significantly affect FR accuracy and fairness. These findings provide practical guidelines for constructing fairer FR systems using synthetic data.
title Investigation of Accuracy and Bias in Face Recognition Trained with Synthetic Data
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2507.20782