Generalizable Audio Spoofing Detection using Non-Semantic Representations

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Das, Arnab, Kheir, Yassine El, Franzreb, Carlos, Herzig, Tim, Polzehl, Tim, Möller, Sebastian
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918132902264832
author Das, Arnab
Kheir, Yassine El
Franzreb, Carlos
Herzig, Tim
Polzehl, Tim
Möller, Sebastian
author_facet Das, Arnab
Kheir, Yassine El
Franzreb, Carlos
Herzig, Tim
Polzehl, Tim
Möller, Sebastian
contents Rapid advancements in generative modeling have made synthetic audio generation easy, making speech-based services vulnerable to spoofing attacks. Consequently, there is a dire need for robust countermeasures more than ever. Existing solutions for deepfake detection are often criticized for lacking generalizability and fail drastically when applied to real-world data. This study proposes a novel method for generalizable spoofing detection leveraging non-semantic universal audio representations. Extensive experiments have been performed to find suitable non-semantic features using TRILL and TRILLsson models. The results indicate that the proposed method achieves comparable performance on the in-domain test set while significantly outperforming state-of-the-art approaches on out-of-domain test sets. Notably, it demonstrates superior generalization on public-domain data, surpassing methods based on hand-crafted features, semantic embeddings, and end-to-end architectures.
format Preprint
id arxiv_https___arxiv_org_abs_2509_00186
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Generalizable Audio Spoofing Detection using Non-Semantic Representations
Das, Arnab
Kheir, Yassine El
Franzreb, Carlos
Herzig, Tim
Polzehl, Tim
Möller, Sebastian
Sound
Artificial Intelligence
Audio and Speech Processing
Rapid advancements in generative modeling have made synthetic audio generation easy, making speech-based services vulnerable to spoofing attacks. Consequently, there is a dire need for robust countermeasures more than ever. Existing solutions for deepfake detection are often criticized for lacking generalizability and fail drastically when applied to real-world data. This study proposes a novel method for generalizable spoofing detection leveraging non-semantic universal audio representations. Extensive experiments have been performed to find suitable non-semantic features using TRILL and TRILLsson models. The results indicate that the proposed method achieves comparable performance on the in-domain test set while significantly outperforming state-of-the-art approaches on out-of-domain test sets. Notably, it demonstrates superior generalization on public-domain data, surpassing methods based on hand-crafted features, semantic embeddings, and end-to-end architectures.
title Generalizable Audio Spoofing Detection using Non-Semantic Representations
topic Sound
Artificial Intelligence
Audio and Speech Processing
url https://arxiv.org/abs/2509.00186