Replay Attacks Against Audio Deepfake Detection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Müller, Nicolas, Kawa, Piotr, Choong, Wei-Herng, Stan, Adriana, Bukkapatnam, Aditya Tirumala, Pizzi, Karla, Wagner, Alexander, Sperl, Philip
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912408149164032
author Müller, Nicolas
Kawa, Piotr
Choong, Wei-Herng
Stan, Adriana
Bukkapatnam, Aditya Tirumala
Pizzi, Karla
Wagner, Alexander
Sperl, Philip
author_facet Müller, Nicolas
Kawa, Piotr
Choong, Wei-Herng
Stan, Adriana
Bukkapatnam, Aditya Tirumala
Pizzi, Karla
Wagner, Alexander
Sperl, Philip
contents We show how replay attacks undermine audio deepfake detection: By playing and re-recording deepfake audio through various speakers and microphones, we make spoofed samples appear authentic to the detection model. To study this phenomenon in more detail, we introduce ReplayDF, a dataset of recordings derived from M-AILABS and MLAAD, featuring 109 speaker-microphone combinations across six languages and four TTS models. It includes diverse acoustic conditions, some highly challenging for detection. Our analysis of six open-source detection models across five datasets reveals significant vulnerability, with the top-performing W2V2-AASIST model's Equal Error Rate (EER) surging from 4.7% to 18.2%. Even with adaptive Room Impulse Response (RIR) retraining, performance remains compromised with an 11.0% EER. We release ReplayDF for non-commercial research use.
format Preprint
id arxiv_https___arxiv_org_abs_2505_14862
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Replay Attacks Against Audio Deepfake Detection
Müller, Nicolas
Kawa, Piotr
Choong, Wei-Herng
Stan, Adriana
Bukkapatnam, Aditya Tirumala
Pizzi, Karla
Wagner, Alexander
Sperl, Philip
Sound
Artificial Intelligence
Audio and Speech Processing
We show how replay attacks undermine audio deepfake detection: By playing and re-recording deepfake audio through various speakers and microphones, we make spoofed samples appear authentic to the detection model. To study this phenomenon in more detail, we introduce ReplayDF, a dataset of recordings derived from M-AILABS and MLAAD, featuring 109 speaker-microphone combinations across six languages and four TTS models. It includes diverse acoustic conditions, some highly challenging for detection. Our analysis of six open-source detection models across five datasets reveals significant vulnerability, with the top-performing W2V2-AASIST model's Equal Error Rate (EER) surging from 4.7% to 18.2%. Even with adaptive Room Impulse Response (RIR) retraining, performance remains compromised with an 11.0% EER. We release ReplayDF for non-commercial research use.
title Replay Attacks Against Audio Deepfake Detection
topic Sound
Artificial Intelligence
Audio and Speech Processing
url https://arxiv.org/abs/2505.14862