FakeSound: Deepfake General Audio Detection

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Xie, Zeyu, Li, Baihan, Xu, Xuenan, Liang, Zheng, Yu, Kai, Wu, Mengyue
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866929382952534016
author Xie, Zeyu
Li, Baihan
Xu, Xuenan
Liang, Zheng
Yu, Kai
Wu, Mengyue
author_facet Xie, Zeyu
Li, Baihan
Xu, Xuenan
Liang, Zheng
Yu, Kai
Wu, Mengyue
contents With the advancement of audio generation, generative models can produce highly realistic audios. However, the proliferation of deepfake general audio can pose negative consequences. Therefore, we propose a new task, deepfake general audio detection, which aims to identify whether audio content is manipulated and to locate deepfake regions. Leveraging an automated manipulation pipeline, a dataset named FakeSound for deepfake general audio detection is proposed, and samples can be viewed on website https://FakeSoundData.github.io. The average binary accuracy of humans on all test sets is consistently below 0.6, which indicates the difficulty humans face in discerning deepfake audio and affirms the efficacy of the FakeSound dataset. A deepfake detection model utilizing a general audio pre-trained model is proposed as a benchmark system. Experimental results demonstrate that the performance of the proposed model surpasses the state-of-the-art in deepfake speech detection and human testers.
format Preprint
id arxiv_https___arxiv_org_abs_2406_08052
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle FakeSound: Deepfake General Audio Detection
Xie, Zeyu
Li, Baihan
Xu, Xuenan
Liang, Zheng
Yu, Kai
Wu, Mengyue
Sound
Audio and Speech Processing
68Txx
I.2
With the advancement of audio generation, generative models can produce highly realistic audios. However, the proliferation of deepfake general audio can pose negative consequences. Therefore, we propose a new task, deepfake general audio detection, which aims to identify whether audio content is manipulated and to locate deepfake regions. Leveraging an automated manipulation pipeline, a dataset named FakeSound for deepfake general audio detection is proposed, and samples can be viewed on website https://FakeSoundData.github.io. The average binary accuracy of humans on all test sets is consistently below 0.6, which indicates the difficulty humans face in discerning deepfake audio and affirms the efficacy of the FakeSound dataset. A deepfake detection model utilizing a general audio pre-trained model is proposed as a benchmark system. Experimental results demonstrate that the performance of the proposed model surpasses the state-of-the-art in deepfake speech detection and human testers.
title FakeSound: Deepfake General Audio Detection
topic Sound
Audio and Speech Processing
68Txx
I.2
url https://arxiv.org/abs/2406.08052