Private kNN-VC: Interpretable Anonymization of Converted Speech

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Franzreb, Carlos, Das, Arnab, Polzehl, Tim, Möller, Sebastian
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913855248007168
author Franzreb, Carlos
Das, Arnab
Polzehl, Tim
Möller, Sebastian
author_facet Franzreb, Carlos
Das, Arnab
Polzehl, Tim
Möller, Sebastian
contents Speaker anonymization seeks to conceal a speaker's identity while preserving the utility of their speech. The achieved privacy is commonly evaluated with a speaker recognition model trained on anonymized speech. Although this represents a strong attack, it is unclear which aspects of speech are exploited to identify the speakers. Our research sets out to unveil these aspects. It starts with kNN-VC, a powerful voice conversion model that performs poorly as an anonymization system, presumably because of prosody leakage. To test this hypothesis, we extend kNN-VC with two interpretable components that anonymize the duration and variation of phones. These components increase privacy significantly, proving that the studied prosodic factors encode speaker identity and are exploited by the privacy attack. Additionally, we show that changes in the target selection algorithm considerably influence the outcome of the privacy attack.
format Preprint
id arxiv_https___arxiv_org_abs_2505_17584
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Private kNN-VC: Interpretable Anonymization of Converted Speech
Franzreb, Carlos
Das, Arnab
Polzehl, Tim
Möller, Sebastian
Audio and Speech Processing
Sound
Speaker anonymization seeks to conceal a speaker's identity while preserving the utility of their speech. The achieved privacy is commonly evaluated with a speaker recognition model trained on anonymized speech. Although this represents a strong attack, it is unclear which aspects of speech are exploited to identify the speakers. Our research sets out to unveil these aspects. It starts with kNN-VC, a powerful voice conversion model that performs poorly as an anonymization system, presumably because of prosody leakage. To test this hypothesis, we extend kNN-VC with two interpretable components that anonymize the duration and variation of phones. These components increase privacy significantly, proving that the studied prosodic factors encode speaker identity and are exploited by the privacy attack. Additionally, we show that changes in the target selection algorithm considerably influence the outcome of the privacy attack.
title Private kNN-VC: Interpretable Anonymization of Converted Speech
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2505.17584