An RFP dataset for Real, Fake, and Partially fake audio detection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: AlAli, Abdulazeez, Theodorakopoulos, George
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910426285998080
author AlAli, Abdulazeez
Theodorakopoulos, George
author_facet AlAli, Abdulazeez
Theodorakopoulos, George
contents Recent advances in deep learning have enabled the creation of natural-sounding synthesised speech. However, attackers have also utilised these tech-nologies to conduct attacks such as phishing. Numerous public datasets have been created to facilitate the development of effective detection models. How-ever, available datasets contain only entirely fake audio; therefore, detection models may miss attacks that replace a short section of the real audio with fake audio. In recognition of this problem, the current paper presents the RFP da-taset, which comprises five distinct audio types: partial fake (PF), audio with noise, voice conversion (VC), text-to-speech (TTS), and real. The data are then used to evaluate several detection models, revealing that the available detec-tion models incur a markedly higher equal error rate (EER) when detecting PF audio instead of entirely fake audio. The lowest EER recorded was 25.42%. Therefore, we believe that creators of detection models must seriously consid-er using datasets like RFP that include PF and other types of fake audio.
format Preprint
id arxiv_https___arxiv_org_abs_2404_17721
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle An RFP dataset for Real, Fake, and Partially fake audio detection
AlAli, Abdulazeez
Theodorakopoulos, George
Sound
Cryptography and Security
Audio and Speech Processing
Recent advances in deep learning have enabled the creation of natural-sounding synthesised speech. However, attackers have also utilised these tech-nologies to conduct attacks such as phishing. Numerous public datasets have been created to facilitate the development of effective detection models. How-ever, available datasets contain only entirely fake audio; therefore, detection models may miss attacks that replace a short section of the real audio with fake audio. In recognition of this problem, the current paper presents the RFP da-taset, which comprises five distinct audio types: partial fake (PF), audio with noise, voice conversion (VC), text-to-speech (TTS), and real. The data are then used to evaluate several detection models, revealing that the available detec-tion models incur a markedly higher equal error rate (EER) when detecting PF audio instead of entirely fake audio. The lowest EER recorded was 25.42%. Therefore, we believe that creators of detection models must seriously consid-er using datasets like RFP that include PF and other types of fake audio.
title An RFP dataset for Real, Fake, and Partially fake audio detection
topic Sound
Cryptography and Security
Audio and Speech Processing
url https://arxiv.org/abs/2404.17721