Post-training for Deepfake Speech Detection

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Ge, Wanying, Wang, Xin, Liu, Xuechen, Yamagishi, Junichi
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914104732549120
author Ge, Wanying
Wang, Xin
Liu, Xuechen
Yamagishi, Junichi
author_facet Ge, Wanying
Wang, Xin
Liu, Xuechen
Yamagishi, Junichi
contents We introduce a post-training approach that adapts self-supervised learning (SSL) models for deepfake speech detection by bridging the gap between general pre-training and domain-specific fine-tuning. We present AntiDeepfake models, a series of post-trained models developed using a large-scale multilingual speech dataset containing over 56,000 hours of genuine speech and 18,000 hours of speech with various artifacts in over one hundred languages. Experimental results show that the post-trained models already exhibit strong robustness and generalization to unseen deepfake speech. When they are further fine-tuned on the Deepfake-Eval-2024 dataset, these models consistently surpass existing state-of-the-art detectors that do not leverage post-training. Model checkpoints and source code are available online.
format Preprint
id arxiv_https___arxiv_org_abs_2506_21090
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Post-training for Deepfake Speech Detection
Ge, Wanying
Wang, Xin
Liu, Xuechen
Yamagishi, Junichi
Audio and Speech Processing
We introduce a post-training approach that adapts self-supervised learning (SSL) models for deepfake speech detection by bridging the gap between general pre-training and domain-specific fine-tuning. We present AntiDeepfake models, a series of post-trained models developed using a large-scale multilingual speech dataset containing over 56,000 hours of genuine speech and 18,000 hours of speech with various artifacts in over one hundred languages. Experimental results show that the post-trained models already exhibit strong robustness and generalization to unseen deepfake speech. When they are further fine-tuned on the Deepfake-Eval-2024 dataset, these models consistently surpass existing state-of-the-art detectors that do not leverage post-training. Model checkpoints and source code are available online.
title Post-training for Deepfake Speech Detection
topic Audio and Speech Processing
url https://arxiv.org/abs/2506.21090