XMAD-Bench: Cross-Domain Multilingual Audio Deepfake Benchmark

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ciobanu, Ioan-Paul, Hiji, Andrei-Iulian, Ristea, Nicolae-Catalin, Irofti, Paul, Rusu, Cristian, Ionescu, Radu Tudor
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915737264717824
author Ciobanu, Ioan-Paul
Hiji, Andrei-Iulian
Ristea, Nicolae-Catalin
Irofti, Paul
Rusu, Cristian
Ionescu, Radu Tudor
author_facet Ciobanu, Ioan-Paul
Hiji, Andrei-Iulian
Ristea, Nicolae-Catalin
Irofti, Paul
Rusu, Cristian
Ionescu, Radu Tudor
contents Recent advances in audio generation led to an increasing number of deepfakes, making the general public more vulnerable to financial scams, identity theft, and misinformation. Audio deepfake detectors promise to alleviate this issue, with many recent studies reporting accuracy rates close to 99%. However, these methods are typically tested in an in-domain setup, where the deepfake samples from the training and test sets are produced by the same generative models. To this end, we introduce XMAD-Bench, a large-scale cross-domain multilingual audio deepfake benchmark comprising 668.8 hours of real and deepfake speech. In our novel dataset, the speakers, the generative methods, and the real audio sources are distinct across training and test splits. This leads to a challenging cross-domain evaluation setup, where audio deepfake detectors can be tested "in the wild". Our in-domain and cross-domain experiments indicate a clear disparity between the in-domain performance of deepfake detectors, which is usually as high as 100%, and the cross-domain performance of the same models, which is sometimes similar to random chance. Our benchmark highlights the need for the development of robust audio deepfake detectors, which maintain their generalization capacity across different languages, speakers, generative methods, and data sources. Our benchmark is publicly released at https://github.com/ristea/xmad-bench/.
format Preprint
id arxiv_https___arxiv_org_abs_2506_00462
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle XMAD-Bench: Cross-Domain Multilingual Audio Deepfake Benchmark
Ciobanu, Ioan-Paul
Hiji, Andrei-Iulian
Ristea, Nicolae-Catalin
Irofti, Paul
Rusu, Cristian
Ionescu, Radu Tudor
Sound
Artificial Intelligence
Computation and Language
Machine Learning
Audio and Speech Processing
Recent advances in audio generation led to an increasing number of deepfakes, making the general public more vulnerable to financial scams, identity theft, and misinformation. Audio deepfake detectors promise to alleviate this issue, with many recent studies reporting accuracy rates close to 99%. However, these methods are typically tested in an in-domain setup, where the deepfake samples from the training and test sets are produced by the same generative models. To this end, we introduce XMAD-Bench, a large-scale cross-domain multilingual audio deepfake benchmark comprising 668.8 hours of real and deepfake speech. In our novel dataset, the speakers, the generative methods, and the real audio sources are distinct across training and test splits. This leads to a challenging cross-domain evaluation setup, where audio deepfake detectors can be tested "in the wild". Our in-domain and cross-domain experiments indicate a clear disparity between the in-domain performance of deepfake detectors, which is usually as high as 100%, and the cross-domain performance of the same models, which is sometimes similar to random chance. Our benchmark highlights the need for the development of robust audio deepfake detectors, which maintain their generalization capacity across different languages, speakers, generative methods, and data sources. Our benchmark is publicly released at https://github.com/ristea/xmad-bench/.
title XMAD-Bench: Cross-Domain Multilingual Audio Deepfake Benchmark
topic Sound
Artificial Intelligence
Computation and Language
Machine Learning
Audio and Speech Processing
url https://arxiv.org/abs/2506.00462