Replay Attacks Against Audio Deepfake Detection
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912408149164032 |
|---|---|
| author | Müller, Nicolas Kawa, Piotr Choong, Wei-Herng Stan, Adriana Bukkapatnam, Aditya Tirumala Pizzi, Karla Wagner, Alexander Sperl, Philip |
| author_facet | Müller, Nicolas Kawa, Piotr Choong, Wei-Herng Stan, Adriana Bukkapatnam, Aditya Tirumala Pizzi, Karla Wagner, Alexander Sperl, Philip |
| contents | We show how replay attacks undermine audio deepfake detection: By playing and re-recording deepfake audio through various speakers and microphones, we make spoofed samples appear authentic to the detection model. To study this phenomenon in more detail, we introduce ReplayDF, a dataset of recordings derived from M-AILABS and MLAAD, featuring 109 speaker-microphone combinations across six languages and four TTS models. It includes diverse acoustic conditions, some highly challenging for detection. Our analysis of six open-source detection models across five datasets reveals significant vulnerability, with the top-performing W2V2-AASIST model's Equal Error Rate (EER) surging from 4.7% to 18.2%. Even with adaptive Room Impulse Response (RIR) retraining, performance remains compromised with an 11.0% EER. We release ReplayDF for non-commercial research use. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2505_14862 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Replay Attacks Against Audio Deepfake Detection Müller, Nicolas Kawa, Piotr Choong, Wei-Herng Stan, Adriana Bukkapatnam, Aditya Tirumala Pizzi, Karla Wagner, Alexander Sperl, Philip Sound Artificial Intelligence Audio and Speech Processing We show how replay attacks undermine audio deepfake detection: By playing and re-recording deepfake audio through various speakers and microphones, we make spoofed samples appear authentic to the detection model. To study this phenomenon in more detail, we introduce ReplayDF, a dataset of recordings derived from M-AILABS and MLAAD, featuring 109 speaker-microphone combinations across six languages and four TTS models. It includes diverse acoustic conditions, some highly challenging for detection. Our analysis of six open-source detection models across five datasets reveals significant vulnerability, with the top-performing W2V2-AASIST model's Equal Error Rate (EER) surging from 4.7% to 18.2%. Even with adaptive Room Impulse Response (RIR) retraining, performance remains compromised with an 11.0% EER. We release ReplayDF for non-commercial research use. |
| title | Replay Attacks Against Audio Deepfake Detection |
| topic | Sound Artificial Intelligence Audio and Speech Processing |
| url | https://arxiv.org/abs/2505.14862 |