Open-Set Source Tracing of Audio Deepfake Systems

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Klein, Nicholas, Tak, Hemlata, Khoury, Elie
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908441673465856
author Klein, Nicholas
Tak, Hemlata
Khoury, Elie
author_facet Klein, Nicholas
Tak, Hemlata
Khoury, Elie
contents Existing research on source tracing of audio deepfake systems has focused primarily on the closed-set scenario, while studies that evaluate open-set performance are limited to a small number of unseen systems. Due to the large number of emerging audio deepfake systems, robust open-set source tracing is critical. We leverage the protocol of the Interspeech 2025 special session on source tracing to evaluate methods for improving open-set source tracing performance. We introduce a novel adaptation to the energy score for out-of-distribution (OOD) detection, softmax energy (SME). We find that replacing the typical temperature-scaled energy score with SME provides a relative average improvement of 31% in the standard FPR95 (false positive rate at true positive rate of 95%) measure. We further explore SME-guided training as well as copy synthesis, codec, and reverberation augmentations, yielding an FPR95 of 8.3%.
format Preprint
id arxiv_https___arxiv_org_abs_2507_06470
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Open-Set Source Tracing of Audio Deepfake Systems
Klein, Nicholas
Tak, Hemlata
Khoury, Elie
Audio and Speech Processing
Sound
Existing research on source tracing of audio deepfake systems has focused primarily on the closed-set scenario, while studies that evaluate open-set performance are limited to a small number of unseen systems. Due to the large number of emerging audio deepfake systems, robust open-set source tracing is critical. We leverage the protocol of the Interspeech 2025 special session on source tracing to evaluate methods for improving open-set source tracing performance. We introduce a novel adaptation to the energy score for out-of-distribution (OOD) detection, softmax energy (SME). We find that replacing the typical temperature-scaled energy score with SME provides a relative average improvement of 31% in the standard FPR95 (false positive rate at true positive rate of 95%) measure. We further explore SME-guided training as well as copy synthesis, codec, and reverberation augmentations, yielding an FPR95 of 8.3%.
title Open-Set Source Tracing of Audio Deepfake Systems
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2507.06470