Massively Annotated Datasets for Assessment of Synthetic and Real Data in Face Recognition

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Neto, Pedro C., Mamede, Rafael M., Albuquerque, Carolina, Gonçalves, Tiago, Sequeira, Ana F.
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911850619207680
author Neto, Pedro C.
Mamede, Rafael M.
Albuquerque, Carolina
Gonçalves, Tiago
Sequeira, Ana F.
author_facet Neto, Pedro C.
Mamede, Rafael M.
Albuquerque, Carolina
Gonçalves, Tiago
Sequeira, Ana F.
contents Face recognition applications have grown in parallel with the size of datasets, complexity of deep learning models and computational power. However, while deep learning models evolve to become more capable and computational power keeps increasing, the datasets available are being retracted and removed from public access. Privacy and ethical concerns are relevant topics within these domains. Through generative artificial intelligence, researchers have put efforts into the development of completely synthetic datasets that can be used to train face recognition systems. Nonetheless, the recent advances have not been sufficient to achieve performance comparable to the state-of-the-art models trained on real data. To study the drift between the performance of models trained on real and synthetic datasets, we leverage a massive attribute classifier (MAC) to create annotations for four datasets: two real and two synthetic. From these annotations, we conduct studies on the distribution of each attribute within all four datasets. Additionally, we further inspect the differences between real and synthetic datasets on the attribute set. When comparing through the Kullback-Leibler divergence we have found differences between real and synthetic samples. Interestingly enough, we have verified that while real samples suffice to explain the synthetic distribution, the opposite could not be further from being true.
format Preprint
id arxiv_https___arxiv_org_abs_2404_15234
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Massively Annotated Datasets for Assessment of Synthetic and Real Data in Face Recognition
Neto, Pedro C.
Mamede, Rafael M.
Albuquerque, Carolina
Gonçalves, Tiago
Sequeira, Ana F.
Computer Vision and Pattern Recognition
Face recognition applications have grown in parallel with the size of datasets, complexity of deep learning models and computational power. However, while deep learning models evolve to become more capable and computational power keeps increasing, the datasets available are being retracted and removed from public access. Privacy and ethical concerns are relevant topics within these domains. Through generative artificial intelligence, researchers have put efforts into the development of completely synthetic datasets that can be used to train face recognition systems. Nonetheless, the recent advances have not been sufficient to achieve performance comparable to the state-of-the-art models trained on real data. To study the drift between the performance of models trained on real and synthetic datasets, we leverage a massive attribute classifier (MAC) to create annotations for four datasets: two real and two synthetic. From these annotations, we conduct studies on the distribution of each attribute within all four datasets. Additionally, we further inspect the differences between real and synthetic datasets on the attribute set. When comparing through the Kullback-Leibler divergence we have found differences between real and synthetic samples. Interestingly enough, we have verified that while real samples suffice to explain the synthetic distribution, the opposite could not be further from being true.
title Massively Annotated Datasets for Assessment of Synthetic and Real Data in Face Recognition
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2404.15234