DFADD: The Diffusion and Flow-Matching Based Audio Deepfake Dataset
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866916392859598848 |
|---|---|
| author | Du, Jiawei Lin, I-Ming Chiu, I-Hsiang Chen, Xuanjun Wu, Haibin Ren, Wenze Tsao, Yu Lee, Hung-yi Jang, Jyh-Shing Roger |
| author_facet | Du, Jiawei Lin, I-Ming Chiu, I-Hsiang Chen, Xuanjun Wu, Haibin Ren, Wenze Tsao, Yu Lee, Hung-yi Jang, Jyh-Shing Roger |
| contents | Mainstream zero-shot TTS production systems like Voicebox and Seed-TTS achieve human parity speech by leveraging Flow-matching and Diffusion models, respectively. Unfortunately, human-level audio synthesis leads to identity misuse and information security issues. Currently, many antispoofing models have been developed against deepfake audio. However, the efficacy of current state-of-the-art anti-spoofing models in countering audio synthesized by diffusion and flowmatching based TTS systems remains unknown. In this paper, we proposed the Diffusion and Flow-matching based Audio Deepfake (DFADD) dataset. The DFADD dataset collected the deepfake audio based on advanced diffusion and flowmatching TTS models. Additionally, we reveal that current anti-spoofing models lack sufficient robustness against highly human-like audio generated by diffusion and flow-matching TTS systems. The proposed DFADD dataset addresses this gap and provides a valuable resource for developing more resilient anti-spoofing models. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2409_08731 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | DFADD: The Diffusion and Flow-Matching Based Audio Deepfake Dataset Du, Jiawei Lin, I-Ming Chiu, I-Hsiang Chen, Xuanjun Wu, Haibin Ren, Wenze Tsao, Yu Lee, Hung-yi Jang, Jyh-Shing Roger Sound Audio and Speech Processing Mainstream zero-shot TTS production systems like Voicebox and Seed-TTS achieve human parity speech by leveraging Flow-matching and Diffusion models, respectively. Unfortunately, human-level audio synthesis leads to identity misuse and information security issues. Currently, many antispoofing models have been developed against deepfake audio. However, the efficacy of current state-of-the-art anti-spoofing models in countering audio synthesized by diffusion and flowmatching based TTS systems remains unknown. In this paper, we proposed the Diffusion and Flow-matching based Audio Deepfake (DFADD) dataset. The DFADD dataset collected the deepfake audio based on advanced diffusion and flowmatching TTS models. Additionally, we reveal that current anti-spoofing models lack sufficient robustness against highly human-like audio generated by diffusion and flow-matching TTS systems. The proposed DFADD dataset addresses this gap and provides a valuable resource for developing more resilient anti-spoofing models. |
| title | DFADD: The Diffusion and Flow-Matching Based Audio Deepfake Dataset |
| topic | Sound Audio and Speech Processing |
| url | https://arxiv.org/abs/2409.08731 |