Scores Know Bobs Voice: Speaker Impersonation Attack

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hwang, Chanwoo, Kim, Sunpill, Tan, Yong Kiam, Liu, Tianchi, Paik, Seunghun, Kim, Dongsoo, Soumik, Mondal, Aung, Khin Mi Mi, Seo, Jae Hong
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910039179001856
author Hwang, Chanwoo
Kim, Sunpill
Tan, Yong Kiam
Liu, Tianchi
Paik, Seunghun
Kim, Dongsoo
Soumik, Mondal
Aung, Khin Mi Mi
Seo, Jae Hong
author_facet Hwang, Chanwoo
Kim, Sunpill
Tan, Yong Kiam
Liu, Tianchi
Paik, Seunghun
Kim, Dongsoo
Soumik, Mondal
Aung, Khin Mi Mi
Seo, Jae Hong
contents Advances in deep learning have enabled the widespread deployment of speaker recognition systems (SRSs), yet they remain vulnerable to score-based impersonation attacks. Existing attacks that operate directly on raw waveforms require a large number of queries due to the difficulty of optimizing in high-dimensional audio spaces. Latent-space optimization within generative models offers improved efficiency, but these latent spaces are shaped by data distribution matching and do not inherently capture speaker-discriminative geometry. As a result, optimization trajectories often fail to align with the adversarial direction needed to maximize victim scores. To address this limitation, we propose an inversion-based generative attack framework that explicitly aligns the latent space of the synthesis model with the discriminative feature space of SRSs. We first analyze the requirements of an inverse model for score-based attacks and introduce a feature-aligned inversion strategy that geometrically synchronizes latent representations with speaker embeddings. This alignment ensures that latent updates directly translate into score improvements. Moreover, it enables new attack paradigms, including subspace-projection-based attacks, which were previously infeasible due to the absence of a faithful feature-to-audio mapping. Experiments show that our method significantly improves query efficiency, achieving competitive attack success rates with on average 10x fewer queries than prior approaches. In particular, the enabled subspace-projection-based attack attains up to 91.65% success using only 50 queries. These findings establish feature-aligned inversion as a key tool for evaluating the robustness of modern SRSs against score-based impersonation threats.
format Preprint
id arxiv_https___arxiv_org_abs_2603_02781
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Scores Know Bobs Voice: Speaker Impersonation Attack
Hwang, Chanwoo
Kim, Sunpill
Tan, Yong Kiam
Liu, Tianchi
Paik, Seunghun
Kim, Dongsoo
Soumik, Mondal
Aung, Khin Mi Mi
Seo, Jae Hong
Cryptography and Security
Artificial Intelligence
Advances in deep learning have enabled the widespread deployment of speaker recognition systems (SRSs), yet they remain vulnerable to score-based impersonation attacks. Existing attacks that operate directly on raw waveforms require a large number of queries due to the difficulty of optimizing in high-dimensional audio spaces. Latent-space optimization within generative models offers improved efficiency, but these latent spaces are shaped by data distribution matching and do not inherently capture speaker-discriminative geometry. As a result, optimization trajectories often fail to align with the adversarial direction needed to maximize victim scores. To address this limitation, we propose an inversion-based generative attack framework that explicitly aligns the latent space of the synthesis model with the discriminative feature space of SRSs. We first analyze the requirements of an inverse model for score-based attacks and introduce a feature-aligned inversion strategy that geometrically synchronizes latent representations with speaker embeddings. This alignment ensures that latent updates directly translate into score improvements. Moreover, it enables new attack paradigms, including subspace-projection-based attacks, which were previously infeasible due to the absence of a faithful feature-to-audio mapping. Experiments show that our method significantly improves query efficiency, achieving competitive attack success rates with on average 10x fewer queries than prior approaches. In particular, the enabled subspace-projection-based attack attains up to 91.65% success using only 50 queries. These findings establish feature-aligned inversion as a key tool for evaluating the robustness of modern SRSs against score-based impersonation threats.
title Scores Know Bobs Voice: Speaker Impersonation Attack
topic Cryptography and Security
Artificial Intelligence
url https://arxiv.org/abs/2603.02781