Impact of Phonetics on Speaker Identity in Adversarial Voice Attack

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Dar, Daniyal Kabir, Yan, Qiben, Xiao, Li, Ross, Arun
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910005212479488
author Dar, Daniyal Kabir
Yan, Qiben
Xiao, Li
Ross, Arun
author_facet Dar, Daniyal Kabir
Yan, Qiben
Xiao, Li
Ross, Arun
contents Adversarial perturbations in speech pose a serious threat to automatic speech recognition (ASR) and speaker verification by introducing subtle waveform modifications that remain imperceptible to humans but can significantly alter system outputs. While targeted attacks on end-to-end ASR models have been widely studied, the phonetic basis of these perturbations and their effect on speaker identity remain underexplored. In this work, we analyze adversarial audio at the phonetic level and show that perturbations exploit systematic confusions such as vowel centralization and consonant substitutions. These distortions not only mislead transcription but also degrade phonetic cues critical for speaker verification, leading to identity drift. Using DeepSpeech as our ASR target, we generate targeted adversarial examples and evaluate their impact on speaker embeddings across genuine and impostor samples. Results across 16 phonetically diverse target phrases demonstrate that adversarial audio induces both transcription errors and identity drift, highlighting the need for phonetic-aware defenses to ensure the robustness of ASR and speaker recognition systems.
format Preprint
id arxiv_https___arxiv_org_abs_2509_15437
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Impact of Phonetics on Speaker Identity in Adversarial Voice Attack
Dar, Daniyal Kabir
Yan, Qiben
Xiao, Li
Ross, Arun
Sound
Artificial Intelligence
Cryptography and Security
Audio and Speech Processing
I.2.0; I.2.7; I.5.4; K.6.5
Adversarial perturbations in speech pose a serious threat to automatic speech recognition (ASR) and speaker verification by introducing subtle waveform modifications that remain imperceptible to humans but can significantly alter system outputs. While targeted attacks on end-to-end ASR models have been widely studied, the phonetic basis of these perturbations and their effect on speaker identity remain underexplored. In this work, we analyze adversarial audio at the phonetic level and show that perturbations exploit systematic confusions such as vowel centralization and consonant substitutions. These distortions not only mislead transcription but also degrade phonetic cues critical for speaker verification, leading to identity drift. Using DeepSpeech as our ASR target, we generate targeted adversarial examples and evaluate their impact on speaker embeddings across genuine and impostor samples. Results across 16 phonetically diverse target phrases demonstrate that adversarial audio induces both transcription errors and identity drift, highlighting the need for phonetic-aware defenses to ensure the robustness of ASR and speaker recognition systems.
title Impact of Phonetics on Speaker Identity in Adversarial Voice Attack
topic Sound
Artificial Intelligence
Cryptography and Security
Audio and Speech Processing
I.2.0; I.2.7; I.5.4; K.6.5
url https://arxiv.org/abs/2509.15437