Multilingual Source Tracing of Speech Deepfakes: A First Benchmark

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xuan, Xi, Xiao, Yang, Das, Rohan Kumar, Kinnunen, Tomi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911093906997248
author Xuan, Xi
Xiao, Yang
Das, Rohan Kumar
Kinnunen, Tomi
author_facet Xuan, Xi
Xiao, Yang
Das, Rohan Kumar
Kinnunen, Tomi
contents Recent progress in generative AI has made it increasingly easy to create natural-sounding deepfake speech from just a few seconds of audio. While these tools support helpful applications, they also raise serious concerns by making it possible to generate convincing fake speech in many languages. Current research has largely focused on detecting fake speech, but little attention has been given to tracing the source models used to generate it. This paper introduces the first benchmark for multilingual speech deepfake source tracing, covering both mono- and cross-lingual scenarios. We comparatively investigate DSP- and SSL-based modeling; examine how SSL representations fine-tuned on different languages impact cross-lingual generalization performance; and evaluate generalization to unseen languages and speakers. Our findings offer the first comprehensive insights into the challenges of identifying speech generation models when training and inference languages differ. The dataset, protocol and code are available at https://github.com/xuanxixi/Multilingual-Source-Tracing.
format Preprint
id arxiv_https___arxiv_org_abs_2508_04143
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Multilingual Source Tracing of Speech Deepfakes: A First Benchmark
Xuan, Xi
Xiao, Yang
Das, Rohan Kumar
Kinnunen, Tomi
Audio and Speech Processing
Computation and Language
Sound
Recent progress in generative AI has made it increasingly easy to create natural-sounding deepfake speech from just a few seconds of audio. While these tools support helpful applications, they also raise serious concerns by making it possible to generate convincing fake speech in many languages. Current research has largely focused on detecting fake speech, but little attention has been given to tracing the source models used to generate it. This paper introduces the first benchmark for multilingual speech deepfake source tracing, covering both mono- and cross-lingual scenarios. We comparatively investigate DSP- and SSL-based modeling; examine how SSL representations fine-tuned on different languages impact cross-lingual generalization performance; and evaluate generalization to unseen languages and speakers. Our findings offer the first comprehensive insights into the challenges of identifying speech generation models when training and inference languages differ. The dataset, protocol and code are available at https://github.com/xuanxixi/Multilingual-Source-Tracing.
title Multilingual Source Tracing of Speech Deepfakes: A First Benchmark
topic Audio and Speech Processing
Computation and Language
Sound
url https://arxiv.org/abs/2508.04143