MDD: a Mask Diffusion Detector to Protect Speaker Verification Systems from Adversarial Perturbations

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Bai, Yibo, Chen, Sizhou, Panariello, Michele, Zhang, Xiao-Lei, Todisco, Massimiliano, Evans, Nicholas
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866909753697894400
author Bai, Yibo
Chen, Sizhou
Panariello, Michele
Zhang, Xiao-Lei
Todisco, Massimiliano
Evans, Nicholas
author_facet Bai, Yibo
Chen, Sizhou
Panariello, Michele
Zhang, Xiao-Lei
Todisco, Massimiliano
Evans, Nicholas
contents Speaker verification systems are increasingly deployed in security-sensitive applications but remain highly vulnerable to adversarial perturbations. In this work, we propose the Mask Diffusion Detector (MDD), a novel adversarial detection and purification framework based on a \textit{text-conditioned masked diffusion model}. During training, MDD applies partial masking to Mel-spectrograms and progressively adds noise through a forward diffusion process, simulating the degradation of clean speech features. A reverse process then reconstructs the clean representation conditioned on the input transcription. Unlike prior approaches, MDD does not require adversarial examples or large-scale pretraining. Experimental results show that MDD achieves strong adversarial detection performance and outperforms prior state-of-the-art methods, including both diffusion-based and neural codec-based approaches. Furthermore, MDD effectively purifies adversarially-manipulated speech, restoring speaker verification performance to levels close to those observed under clean conditions. These findings demonstrate the potential of diffusion-based masking strategies for secure and reliable speaker verification systems.
format Preprint
id arxiv_https___arxiv_org_abs_2508_19180
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MDD: a Mask Diffusion Detector to Protect Speaker Verification Systems from Adversarial Perturbations
Bai, Yibo
Chen, Sizhou
Panariello, Michele
Zhang, Xiao-Lei
Todisco, Massimiliano
Evans, Nicholas
Audio and Speech Processing
Sound
Speaker verification systems are increasingly deployed in security-sensitive applications but remain highly vulnerable to adversarial perturbations. In this work, we propose the Mask Diffusion Detector (MDD), a novel adversarial detection and purification framework based on a \textit{text-conditioned masked diffusion model}. During training, MDD applies partial masking to Mel-spectrograms and progressively adds noise through a forward diffusion process, simulating the degradation of clean speech features. A reverse process then reconstructs the clean representation conditioned on the input transcription. Unlike prior approaches, MDD does not require adversarial examples or large-scale pretraining. Experimental results show that MDD achieves strong adversarial detection performance and outperforms prior state-of-the-art methods, including both diffusion-based and neural codec-based approaches. Furthermore, MDD effectively purifies adversarially-manipulated speech, restoring speaker verification performance to levels close to those observed under clean conditions. These findings demonstrate the potential of diffusion-based masking strategies for secure and reliable speaker verification systems.
title MDD: a Mask Diffusion Detector to Protect Speaker Verification Systems from Adversarial Perturbations
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2508.19180