Unveiling Synthetic Faces: How Synthetic Datasets Can Expose Real Identities

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shahreza, Hatef Otroshi, Marcel, Sébastien
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910679365058560
author Shahreza, Hatef Otroshi
Marcel, Sébastien
author_facet Shahreza, Hatef Otroshi
Marcel, Sébastien
contents Synthetic data generation is gaining increasing popularity in different computer vision applications. Existing state-of-the-art face recognition models are trained using large-scale face datasets, which are crawled from the Internet and raise privacy and ethical concerns. To address such concerns, several works have proposed generating synthetic face datasets to train face recognition models. However, these methods depend on generative models, which are trained on real face images. In this work, we design a simple yet effective membership inference attack to systematically study if any of the existing synthetic face recognition datasets leak any information from the real data used to train the generator model. We provide an extensive study on 6 state-of-the-art synthetic face recognition datasets, and show that in all these synthetic datasets, several samples from the original real dataset are leaked. To our knowledge, this paper is the first work which shows the leakage from training data of generator models into the generated synthetic face recognition datasets. Our study demonstrates privacy pitfalls in synthetic face recognition datasets and paves the way for future studies on generating responsible synthetic face datasets.
format Preprint
id arxiv_https___arxiv_org_abs_2410_24015
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Unveiling Synthetic Faces: How Synthetic Datasets Can Expose Real Identities
Shahreza, Hatef Otroshi
Marcel, Sébastien
Computer Vision and Pattern Recognition
Synthetic data generation is gaining increasing popularity in different computer vision applications. Existing state-of-the-art face recognition models are trained using large-scale face datasets, which are crawled from the Internet and raise privacy and ethical concerns. To address such concerns, several works have proposed generating synthetic face datasets to train face recognition models. However, these methods depend on generative models, which are trained on real face images. In this work, we design a simple yet effective membership inference attack to systematically study if any of the existing synthetic face recognition datasets leak any information from the real data used to train the generator model. We provide an extensive study on 6 state-of-the-art synthetic face recognition datasets, and show that in all these synthetic datasets, several samples from the original real dataset are leaked. To our knowledge, this paper is the first work which shows the leakage from training data of generator models into the generated synthetic face recognition datasets. Our study demonstrates privacy pitfalls in synthetic face recognition datasets and paves the way for future studies on generating responsible synthetic face datasets.
title Unveiling Synthetic Faces: How Synthetic Datasets Can Expose Real Identities
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2410.24015