DiffAttack: Diffusion-based Timbre-reserved Adversarial Attack in Speaker Identification

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wang, Qing, Yao, Jixun, Sun, Zhaokai, Guo, Pengcheng, Xie, Lei, Hansen, John H. L.
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913642670194688
author Wang, Qing
Yao, Jixun
Sun, Zhaokai
Guo, Pengcheng
Xie, Lei
Hansen, John H. L.
author_facet Wang, Qing
Yao, Jixun
Sun, Zhaokai
Guo, Pengcheng
Xie, Lei
Hansen, John H. L.
contents Being a form of biometric identification, the security of the speaker identification (SID) system is of utmost importance. To better understand the robustness of SID systems, we aim to perform more realistic attacks in SID, which are challenging for both humans and machines to detect. In this study, we propose DiffAttack, a novel timbre-reserved adversarial attack approach that exploits the capability of a diffusion-based voice conversion (DiffVC) model to generate adversarial fake audio with distinct target speaker attribution. By introducing adversarial constraints into the generative process of the diffusion-based voice conversion model, we craft fake samples that effectively mislead target models while preserving speaker-wise characteristics. Specifically, inspired by the use of randomly sampled Gaussian noise in conventional adversarial attacks and diffusion processes, we incorporate adversarial constraints into the reverse diffusion process. These constraints subtly guide the reverse diffusion process toward aligning with the target speaker distribution. Our experiments on the LibriTTS dataset indicate that DiffAttack significantly improves the attack success rate compared to vanilla DiffVC and other methods. Moreover, objective and subjective evaluations demonstrate that introducing adversarial constraints does not compromise the speech quality generated by the DiffVC model.
format Preprint
id arxiv_https___arxiv_org_abs_2501_05127
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DiffAttack: Diffusion-based Timbre-reserved Adversarial Attack in Speaker Identification
Wang, Qing
Yao, Jixun
Sun, Zhaokai
Guo, Pengcheng
Xie, Lei
Hansen, John H. L.
Sound
Audio and Speech Processing
Being a form of biometric identification, the security of the speaker identification (SID) system is of utmost importance. To better understand the robustness of SID systems, we aim to perform more realistic attacks in SID, which are challenging for both humans and machines to detect. In this study, we propose DiffAttack, a novel timbre-reserved adversarial attack approach that exploits the capability of a diffusion-based voice conversion (DiffVC) model to generate adversarial fake audio with distinct target speaker attribution. By introducing adversarial constraints into the generative process of the diffusion-based voice conversion model, we craft fake samples that effectively mislead target models while preserving speaker-wise characteristics. Specifically, inspired by the use of randomly sampled Gaussian noise in conventional adversarial attacks and diffusion processes, we incorporate adversarial constraints into the reverse diffusion process. These constraints subtly guide the reverse diffusion process toward aligning with the target speaker distribution. Our experiments on the LibriTTS dataset indicate that DiffAttack significantly improves the attack success rate compared to vanilla DiffVC and other methods. Moreover, objective and subjective evaluations demonstrate that introducing adversarial constraints does not compromise the speech quality generated by the DiffVC model.
title DiffAttack: Diffusion-based Timbre-reserved Adversarial Attack in Speaker Identification
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2501.05127