Unmasking real-world audio deepfakes: A data-centric approach

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Combei, David, Stan, Adriana, Oneata, Dan, Müller, Nicolas, Cucu, Horia
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866911183313829888
author Combei, David
Stan, Adriana
Oneata, Dan
Müller, Nicolas
Cucu, Horia
author_facet Combei, David
Stan, Adriana
Oneata, Dan
Müller, Nicolas
Cucu, Horia
contents The growing prevalence of real-world deepfakes presents a critical challenge for existing detection systems, which are often evaluated on datasets collected just for scientific purposes. To address this gap, we introduce a novel dataset of real-world audio deepfakes. Our analysis reveals that these real-world examples pose significant challenges, even for the most performant detection models. Rather than increasing model complexity or exhaustively search for a better alternative, in this work we focus on a data-centric paradigm, employing strategies like dataset curation, pruning, and augmentation to improve model robustness and generalization. Through these methods, we achieve a 55% relative reduction in EER on the In-the-Wild dataset, reaching an absolute EER of 1.7%, and a 63% reduction on our newly proposed real-world deepfakes dataset, AI4T. These results highlight the transformative potential of data-centric approaches in enhancing deepfake detection for real-world applications. Code and data available at: https://github.com/davidcombei/AI4T.
format Preprint
id arxiv_https___arxiv_org_abs_2506_09606
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Unmasking real-world audio deepfakes: A data-centric approach
Combei, David
Stan, Adriana
Oneata, Dan
Müller, Nicolas
Cucu, Horia
Audio and Speech Processing
The growing prevalence of real-world deepfakes presents a critical challenge for existing detection systems, which are often evaluated on datasets collected just for scientific purposes. To address this gap, we introduce a novel dataset of real-world audio deepfakes. Our analysis reveals that these real-world examples pose significant challenges, even for the most performant detection models. Rather than increasing model complexity or exhaustively search for a better alternative, in this work we focus on a data-centric paradigm, employing strategies like dataset curation, pruning, and augmentation to improve model robustness and generalization. Through these methods, we achieve a 55% relative reduction in EER on the In-the-Wild dataset, reaching an absolute EER of 1.7%, and a 63% reduction on our newly proposed real-world deepfakes dataset, AI4T. These results highlight the transformative potential of data-centric approaches in enhancing deepfake detection for real-world applications. Code and data available at: https://github.com/davidcombei/AI4T.
title Unmasking real-world audio deepfakes: A data-centric approach
topic Audio and Speech Processing
url https://arxiv.org/abs/2506.09606