A Comprehensive Survey with Critical Analysis for Deepfake Speech Detection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Pham, Lam, Lam, Phat, Tran, Dat, Tang, Hieu, Nguyen, Tin, Schindler, Alexander, Skopik, Florian, Polonsky, Alexander, Vu, Canh
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915212644319232
author Pham, Lam
Lam, Phat
Tran, Dat
Tang, Hieu
Nguyen, Tin
Schindler, Alexander
Skopik, Florian
Polonsky, Alexander
Vu, Canh
author_facet Pham, Lam
Lam, Phat
Tran, Dat
Tang, Hieu
Nguyen, Tin
Schindler, Alexander
Skopik, Florian
Polonsky, Alexander
Vu, Canh
contents Thanks to advancements in deep learning, speech generation systems now power a variety of real-world applications, such as text-to-speech for individuals with speech disorders, voice chatbots in call centers, cross-linguistic speech translation, etc. While these systems can autonomously generate human-like speech and replicate specific voices, they also pose risks when misused for malicious purposes. This motivates the research community to develop models for detecting synthesized speech (e.g., fake speech) generated by deep-learning-based models, referred to as the Deepfake Speech Detection task. As the Deepfake Speech Detection task has emerged in recent years, there are not many survey papers proposed for this task. Additionally, existing surveys for the Deepfake Speech Detection task tend to summarize techniques used to construct a Deepfake Speech Detection system rather than providing a thorough analysis. This gap motivated us to conduct a comprehensive survey, providing a critical analysis of the challenges and developments in Deepfake Speech Detection. Our survey is innovatively structured, offering an in-depth analysis of current challenge competitions, public datasets, and the deep-learning techniques that provide enhanced solutions to address existing challenges in the field. From our analysis, we propose hypotheses on leveraging and combining specific deep learning techniques to improve the effectiveness of Deepfake Speech Detection systems. Beyond conducting a survey, we perform extensive experiments to validate these hypotheses and propose a highly competitive model for the task of Deepfake Speech Detection. Given the analysis and the experimental results, we finally indicate potential and promising research directions for the Deepfake Speech Detection task.
format Preprint
id arxiv_https___arxiv_org_abs_2409_15180
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle A Comprehensive Survey with Critical Analysis for Deepfake Speech Detection
Pham, Lam
Lam, Phat
Tran, Dat
Tang, Hieu
Nguyen, Tin
Schindler, Alexander
Skopik, Florian
Polonsky, Alexander
Vu, Canh
Sound
Audio and Speech Processing
Thanks to advancements in deep learning, speech generation systems now power a variety of real-world applications, such as text-to-speech for individuals with speech disorders, voice chatbots in call centers, cross-linguistic speech translation, etc. While these systems can autonomously generate human-like speech and replicate specific voices, they also pose risks when misused for malicious purposes. This motivates the research community to develop models for detecting synthesized speech (e.g., fake speech) generated by deep-learning-based models, referred to as the Deepfake Speech Detection task. As the Deepfake Speech Detection task has emerged in recent years, there are not many survey papers proposed for this task. Additionally, existing surveys for the Deepfake Speech Detection task tend to summarize techniques used to construct a Deepfake Speech Detection system rather than providing a thorough analysis. This gap motivated us to conduct a comprehensive survey, providing a critical analysis of the challenges and developments in Deepfake Speech Detection. Our survey is innovatively structured, offering an in-depth analysis of current challenge competitions, public datasets, and the deep-learning techniques that provide enhanced solutions to address existing challenges in the field. From our analysis, we propose hypotheses on leveraging and combining specific deep learning techniques to improve the effectiveness of Deepfake Speech Detection systems. Beyond conducting a survey, we perform extensive experiments to validate these hypotheses and propose a highly competitive model for the task of Deepfake Speech Detection. Given the analysis and the experimental results, we finally indicate potential and promising research directions for the Deepfake Speech Detection task.
title A Comprehensive Survey with Critical Analysis for Deepfake Speech Detection
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2409.15180