Are audio DeepFake detection models polyglots?

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Marek, Bartłomiej, Kawa, Piotr, Syga, Piotr
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915431257735168
author Marek, Bartłomiej
Kawa, Piotr
Syga, Piotr
author_facet Marek, Bartłomiej
Kawa, Piotr
Syga, Piotr
contents Since the majority of audio DeepFake (DF) detection methods are trained on English-centric datasets, their applicability to non-English languages remains largely unexplored. In this work, we present a benchmark for the multilingual audio DF detection challenge by evaluating various adaptation strategies. Our experiments focus on analyzing models trained on English benchmark datasets, as well as intra-linguistic (same-language) and cross-linguistic adaptation approaches. Our results indicate considerable variations in detection efficacy, highlighting the difficulties of multilingual settings. We show that limiting the dataset to English negatively impacts the efficacy, while stressing the importance of the data in the target language.
format Preprint
id arxiv_https___arxiv_org_abs_2412_17924
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Are audio DeepFake detection models polyglots?
Marek, Bartłomiej
Kawa, Piotr
Syga, Piotr
Sound
Audio and Speech Processing
Since the majority of audio DeepFake (DF) detection methods are trained on English-centric datasets, their applicability to non-English languages remains largely unexplored. In this work, we present a benchmark for the multilingual audio DF detection challenge by evaluating various adaptation strategies. Our experiments focus on analyzing models trained on English benchmark datasets, as well as intra-linguistic (same-language) and cross-linguistic adaptation approaches. Our results indicate considerable variations in detection efficacy, highlighting the difficulties of multilingual settings. We show that limiting the dataset to English negatively impacts the efficacy, while stressing the importance of the data in the target language.
title Are audio DeepFake detection models polyglots?
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2412.17924