Cross-Domain Audio Deepfake Detection: Dataset and Analysis

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Li, Yuang, Zhang, Min, Ren, Mengxin, Ma, Miaomiao, Wei, Daimeng, Yang, Hao
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909320180924416
author Li, Yuang
Zhang, Min
Ren, Mengxin
Ma, Miaomiao
Wei, Daimeng
Yang, Hao
author_facet Li, Yuang
Zhang, Min
Ren, Mengxin
Ma, Miaomiao
Wei, Daimeng
Yang, Hao
contents Audio deepfake detection (ADD) is essential for preventing the misuse of synthetic voices that may infringe on personal rights and privacy. Recent zero-shot text-to-speech (TTS) models pose higher risks as they can clone voices with a single utterance. However, the existing ADD datasets are outdated, leading to suboptimal generalization of detection models. In this paper, we construct a new cross-domain ADD dataset comprising over 300 hours of speech data that is generated by five advanced zero-shot TTS models. To simulate real-world scenarios, we employ diverse attack methods and audio prompts from different datasets. Experiments show that, through novel attack-augmented training, the Wav2Vec2-large and Whisper-medium models achieve equal error rates of 4.1\% and 6.5\% respectively. Additionally, we demonstrate our models' outstanding few-shot ADD ability by fine-tuning with just one minute of target-domain data. Nonetheless, neural codec compressors greatly affect the detection accuracy, necessitating further research.
format Preprint
id arxiv_https___arxiv_org_abs_2404_04904
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Cross-Domain Audio Deepfake Detection: Dataset and Analysis
Li, Yuang
Zhang, Min
Ren, Mengxin
Ma, Miaomiao
Wei, Daimeng
Yang, Hao
Sound
Artificial Intelligence
Audio and Speech Processing
Audio deepfake detection (ADD) is essential for preventing the misuse of synthetic voices that may infringe on personal rights and privacy. Recent zero-shot text-to-speech (TTS) models pose higher risks as they can clone voices with a single utterance. However, the existing ADD datasets are outdated, leading to suboptimal generalization of detection models. In this paper, we construct a new cross-domain ADD dataset comprising over 300 hours of speech data that is generated by five advanced zero-shot TTS models. To simulate real-world scenarios, we employ diverse attack methods and audio prompts from different datasets. Experiments show that, through novel attack-augmented training, the Wav2Vec2-large and Whisper-medium models achieve equal error rates of 4.1\% and 6.5\% respectively. Additionally, we demonstrate our models' outstanding few-shot ADD ability by fine-tuning with just one minute of target-domain data. Nonetheless, neural codec compressors greatly affect the detection accuracy, necessitating further research.
title Cross-Domain Audio Deepfake Detection: Dataset and Analysis
topic Sound
Artificial Intelligence
Audio and Speech Processing
url https://arxiv.org/abs/2404.04904