Generalizable Audio Deepfake Detection via Latent Space Refinement and Augmentation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Huang, Wen, Gu, Yanmei, Wang, Zhiming, Zhu, Huijia, Qian, Yanmin
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917901508804608
author Huang, Wen
Gu, Yanmei
Wang, Zhiming
Zhu, Huijia
Qian, Yanmin
author_facet Huang, Wen
Gu, Yanmei
Wang, Zhiming
Zhu, Huijia
Qian, Yanmin
contents Advances in speech synthesis technologies, like text-to-speech (TTS) and voice conversion (VC), have made detecting deepfake speech increasingly challenging. Spoofing countermeasures often struggle to generalize effectively, particularly when faced with unseen attacks. To address this, we propose a novel strategy that integrates Latent Space Refinement (LSR) and Latent Space Augmentation (LSA) to improve the generalization of deepfake detection systems. LSR introduces multiple learnable prototypes for the spoof class, refining the latent space to better capture the intricate variations within spoofed data. LSA further diversifies spoofed data representations by applying augmentation techniques directly in the latent space, enabling the model to learn a broader range of spoofing patterns. We evaluated our approach on four representative datasets, i.e. ASVspoof 2019 LA, ASVspoof 2021 LA and DF, and In-The-Wild. The results show that LSR and LSA perform well individually, and their integration achieves competitive results, matching or surpassing current state-of-the-art methods.
format Preprint
id arxiv_https___arxiv_org_abs_2501_14240
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Generalizable Audio Deepfake Detection via Latent Space Refinement and Augmentation
Huang, Wen
Gu, Yanmei
Wang, Zhiming
Zhu, Huijia
Qian, Yanmin
Audio and Speech Processing
Sound
Advances in speech synthesis technologies, like text-to-speech (TTS) and voice conversion (VC), have made detecting deepfake speech increasingly challenging. Spoofing countermeasures often struggle to generalize effectively, particularly when faced with unseen attacks. To address this, we propose a novel strategy that integrates Latent Space Refinement (LSR) and Latent Space Augmentation (LSA) to improve the generalization of deepfake detection systems. LSR introduces multiple learnable prototypes for the spoof class, refining the latent space to better capture the intricate variations within spoofed data. LSA further diversifies spoofed data representations by applying augmentation techniques directly in the latent space, enabling the model to learn a broader range of spoofing patterns. We evaluated our approach on four representative datasets, i.e. ASVspoof 2019 LA, ASVspoof 2021 LA and DF, and In-The-Wild. The results show that LSR and LSA perform well individually, and their integration achieves competitive results, matching or surpassing current state-of-the-art methods.
title Generalizable Audio Deepfake Detection via Latent Space Refinement and Augmentation
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2501.14240