MDD: a Mask Diffusion Detector to Protect Speaker Verification Systems from Adversarial Perturbations
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | , , , , , |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866909753697894400 |
|---|---|
| author | Bai, Yibo Chen, Sizhou Panariello, Michele Zhang, Xiao-Lei Todisco, Massimiliano Evans, Nicholas |
| author_facet | Bai, Yibo Chen, Sizhou Panariello, Michele Zhang, Xiao-Lei Todisco, Massimiliano Evans, Nicholas |
| contents | Speaker verification systems are increasingly deployed in security-sensitive applications but remain highly vulnerable to adversarial perturbations. In this work, we propose the Mask Diffusion Detector (MDD), a novel adversarial detection and purification framework based on a \textit{text-conditioned masked diffusion model}. During training, MDD applies partial masking to Mel-spectrograms and progressively adds noise through a forward diffusion process, simulating the degradation of clean speech features. A reverse process then reconstructs the clean representation conditioned on the input transcription. Unlike prior approaches, MDD does not require adversarial examples or large-scale pretraining. Experimental results show that MDD achieves strong adversarial detection performance and outperforms prior state-of-the-art methods, including both diffusion-based and neural codec-based approaches. Furthermore, MDD effectively purifies adversarially-manipulated speech, restoring speaker verification performance to levels close to those observed under clean conditions. These findings demonstrate the potential of diffusion-based masking strategies for secure and reliable speaker verification systems. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2508_19180 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | MDD: a Mask Diffusion Detector to Protect Speaker Verification Systems from Adversarial Perturbations Bai, Yibo Chen, Sizhou Panariello, Michele Zhang, Xiao-Lei Todisco, Massimiliano Evans, Nicholas Audio and Speech Processing Sound Speaker verification systems are increasingly deployed in security-sensitive applications but remain highly vulnerable to adversarial perturbations. In this work, we propose the Mask Diffusion Detector (MDD), a novel adversarial detection and purification framework based on a \textit{text-conditioned masked diffusion model}. During training, MDD applies partial masking to Mel-spectrograms and progressively adds noise through a forward diffusion process, simulating the degradation of clean speech features. A reverse process then reconstructs the clean representation conditioned on the input transcription. Unlike prior approaches, MDD does not require adversarial examples or large-scale pretraining. Experimental results show that MDD achieves strong adversarial detection performance and outperforms prior state-of-the-art methods, including both diffusion-based and neural codec-based approaches. Furthermore, MDD effectively purifies adversarially-manipulated speech, restoring speaker verification performance to levels close to those observed under clean conditions. These findings demonstrate the potential of diffusion-based masking strategies for secure and reliable speaker verification systems. |
| title | MDD: a Mask Diffusion Detector to Protect Speaker Verification Systems from Adversarial Perturbations |
| topic | Audio and Speech Processing Sound |
| url | https://arxiv.org/abs/2508.19180 |