A Survey on Speech Deepfake Detection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Menglu, Ahmadiadli, Yasaman, Zhang, Xiao-Ping
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915389400678400
author Li, Menglu
Ahmadiadli, Yasaman
Zhang, Xiao-Ping
author_facet Li, Menglu
Ahmadiadli, Yasaman
Zhang, Xiao-Ping
contents The availability of smart devices leads to an exponential increase in multimedia content. However, advancements in deep learning have also enabled the creation of highly sophisticated Deepfake content, including speech Deepfakes, which pose a serious threat by generating realistic voices and spreading misinformation. To combat this, numerous challenges have been organized to advance speech Deepfake detection techniques. In this survey, we systematically analyze more than 200 papers published up to March 2024. We provide a comprehensive review of each component in the detection pipeline, including model architectures, optimization techniques, generalizability, evaluation metrics, performance comparisons, available datasets, and open source availability. For each aspect, we assess recent progress and discuss ongoing challenges. In addition, we explore emerging topics such as partial Deepfake detection, cross-dataset evaluation, and defences against adversarial attacks, while suggesting promising research directions. This survey not only identifies the current state of the art to establish strong baselines for future experiments but also offers clear guidance for researchers aiming to enhance speech Deepfake detection systems.
format Preprint
id arxiv_https___arxiv_org_abs_2404_13914
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle A Survey on Speech Deepfake Detection
Li, Menglu
Ahmadiadli, Yasaman
Zhang, Xiao-Ping
Sound
Cryptography and Security
Multimedia
Audio and Speech Processing
The availability of smart devices leads to an exponential increase in multimedia content. However, advancements in deep learning have also enabled the creation of highly sophisticated Deepfake content, including speech Deepfakes, which pose a serious threat by generating realistic voices and spreading misinformation. To combat this, numerous challenges have been organized to advance speech Deepfake detection techniques. In this survey, we systematically analyze more than 200 papers published up to March 2024. We provide a comprehensive review of each component in the detection pipeline, including model architectures, optimization techniques, generalizability, evaluation metrics, performance comparisons, available datasets, and open source availability. For each aspect, we assess recent progress and discuss ongoing challenges. In addition, we explore emerging topics such as partial Deepfake detection, cross-dataset evaluation, and defences against adversarial attacks, while suggesting promising research directions. This survey not only identifies the current state of the art to establish strong baselines for future experiments but also offers clear guidance for researchers aiming to enhance speech Deepfake detection systems.
title A Survey on Speech Deepfake Detection
topic Sound
Cryptography and Security
Multimedia
Audio and Speech Processing
url https://arxiv.org/abs/2404.13914