Saved in:
Bibliographic Details
Main Authors: Gaznepoglu, Ünal Ege, Leschanowsky, Anna, Aloradi, Ahmad, Singh, Prachi, Tenbrinck, Daniel, Habets, Emanuël A. P., Peters, Nils
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2506.09521
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918514043912192
author Gaznepoglu, Ünal Ege
Leschanowsky, Anna
Aloradi, Ahmad
Singh, Prachi
Tenbrinck, Daniel
Habets, Emanuël A. P.
Peters, Nils
author_facet Gaznepoglu, Ünal Ege
Leschanowsky, Anna
Aloradi, Ahmad
Singh, Prachi
Tenbrinck, Daniel
Habets, Emanuël A. P.
Peters, Nils
contents Speaker anonymization systems hide the identity of speakers while preserving other information such as linguistic content and emotions. To evaluate their privacy benefits, attacks in the form of automatic speaker verification (ASV) systems are employed. In this study, we assess the impact of intra-speaker linguistic content similarity in the attacker training and evaluation datasets, by adapting BERT, a language model, as an ASV system. On the VoicePrivacy Attacker Challenge datasets, our method achieves a mean equal error rate (EER) of 35%, with certain speakers attaining EERs as low as 2%, based solely on the textual content of their utterances. Our explainability study reveals that the system decisions are linked to semantically similar keywords within utterances, stemming from how LibriSpeech is curated. Our study suggests reworking the VoicePrivacy datasets to ensure a fair and unbiased evaluation and challenge the reliance on global EER for privacy evaluations.
format Preprint
id arxiv_https___arxiv_org_abs_2506_09521
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle You Are What You Say: Exploiting Linguistic Content for VoicePrivacy Attacks
Gaznepoglu, Ünal Ege
Leschanowsky, Anna
Aloradi, Ahmad
Singh, Prachi
Tenbrinck, Daniel
Habets, Emanuël A. P.
Peters, Nils
Audio and Speech Processing
Computation and Language
Speaker anonymization systems hide the identity of speakers while preserving other information such as linguistic content and emotions. To evaluate their privacy benefits, attacks in the form of automatic speaker verification (ASV) systems are employed. In this study, we assess the impact of intra-speaker linguistic content similarity in the attacker training and evaluation datasets, by adapting BERT, a language model, as an ASV system. On the VoicePrivacy Attacker Challenge datasets, our method achieves a mean equal error rate (EER) of 35%, with certain speakers attaining EERs as low as 2%, based solely on the textual content of their utterances. Our explainability study reveals that the system decisions are linked to semantically similar keywords within utterances, stemming from how LibriSpeech is curated. Our study suggests reworking the VoicePrivacy datasets to ensure a fair and unbiased evaluation and challenge the reliance on global EER for privacy evaluations.
title You Are What You Say: Exploiting Linguistic Content for VoicePrivacy Attacks
topic Audio and Speech Processing
Computation and Language
url https://arxiv.org/abs/2506.09521