Privacy-oriented manipulation of speaker representations

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Teixeira, Francisco, Abad, Alberto, Raj, Bhiksha, Trancoso, Isabel
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912021545484288
author Teixeira, Francisco
Abad, Alberto
Raj, Bhiksha
Trancoso, Isabel
author_facet Teixeira, Francisco
Abad, Alberto
Raj, Bhiksha
Trancoso, Isabel
contents Speaker embeddings are ubiquitous, with applications ranging from speaker recognition and diarization to speech synthesis and voice anonymisation. The amount of information held by these embeddings lends them versatility, but also raises privacy concerns. Speaker embeddings have been shown to contain information on age, sex, health and more, which speakers may want to keep private, especially when this information is not required for the target task. In this work, we propose a method for removing and manipulating private attributes from speaker embeddings that leverages a Vector-Quantized Variational Autoencoder architecture, combined with an adversarial classifier and a novel mutual information loss. We validate our model on two attributes, sex and age, and perform experiments with ignorant and fully-informed attackers, and with in-domain and out-of-domain data.
format Preprint
id arxiv_https___arxiv_org_abs_2310_06652
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Privacy-oriented manipulation of speaker representations
Teixeira, Francisco
Abad, Alberto
Raj, Bhiksha
Trancoso, Isabel
Audio and Speech Processing
Speaker embeddings are ubiquitous, with applications ranging from speaker recognition and diarization to speech synthesis and voice anonymisation. The amount of information held by these embeddings lends them versatility, but also raises privacy concerns. Speaker embeddings have been shown to contain information on age, sex, health and more, which speakers may want to keep private, especially when this information is not required for the target task. In this work, we propose a method for removing and manipulating private attributes from speaker embeddings that leverages a Vector-Quantized Variational Autoencoder architecture, combined with an adversarial classifier and a novel mutual information loss. We validate our model on two attributes, sex and age, and perform experiments with ignorant and fully-informed attackers, and with in-domain and out-of-domain data.
title Privacy-oriented manipulation of speaker representations
topic Audio and Speech Processing
url https://arxiv.org/abs/2310.06652