Universal Semantic Disentangled Privacy-preserving Speech Representation Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Vecino, Biel Tura, Maji, Subhadeep, Varier, Aravind, Bonafonte, Antonio, Valles, Ivan, Owen, Michael, Rädel, Leif, Strimel, Grant, Feyisetan, Seyi, Chicote, Roberto Barra, Rastrow, Ariya, Papayiannis, Constantinos, Leutnant, Volker, Wood, Trevor
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909617403985920
author Vecino, Biel Tura
Maji, Subhadeep
Varier, Aravind
Bonafonte, Antonio
Valles, Ivan
Owen, Michael
Rädel, Leif
Strimel, Grant
Feyisetan, Seyi
Chicote, Roberto Barra
Rastrow, Ariya
Papayiannis, Constantinos
Leutnant, Volker
Wood, Trevor
author_facet Vecino, Biel Tura
Maji, Subhadeep
Varier, Aravind
Bonafonte, Antonio
Valles, Ivan
Owen, Michael
Rädel, Leif
Strimel, Grant
Feyisetan, Seyi
Chicote, Roberto Barra
Rastrow, Ariya
Papayiannis, Constantinos
Leutnant, Volker
Wood, Trevor
contents The use of audio recordings of human speech to train LLMs poses privacy concerns due to these models' potential to generate outputs that closely resemble artifacts in the training data. In this study, we propose a speaker privacy-preserving representation learning method through the Universal Speech Codec (USC), a computationally efficient encoder-decoder model that disentangles speech into: (i) privacy-preserving semantically rich representations, capturing content and speech paralinguistics, and (ii) residual acoustic and speaker representations that enables high-fidelity reconstruction. Extensive evaluations presented show that USC's semantic representation preserves content, prosody, and sentiment, while removing potentially identifiable speaker attributes. Combining both representations, USC achieves state-of-the-art speech reconstruction. Additionally, we introduce an evaluation methodology for measuring privacy-preserving properties, aligning with perceptual tests. We compare USC against other codecs in the literature and demonstrate its effectiveness on privacy-preserving representation learning, illustrating the trade-offs of speaker anonymization, paralinguistics retention and content preservation in the learned semantic representations. Audio samples are shared in https://www.amazon.science/usc-samples.
format Preprint
id arxiv_https___arxiv_org_abs_2505_13085
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Universal Semantic Disentangled Privacy-preserving Speech Representation Learning
Vecino, Biel Tura
Maji, Subhadeep
Varier, Aravind
Bonafonte, Antonio
Valles, Ivan
Owen, Michael
Rädel, Leif
Strimel, Grant
Feyisetan, Seyi
Chicote, Roberto Barra
Rastrow, Ariya
Papayiannis, Constantinos
Leutnant, Volker
Wood, Trevor
Audio and Speech Processing
Machine Learning
The use of audio recordings of human speech to train LLMs poses privacy concerns due to these models' potential to generate outputs that closely resemble artifacts in the training data. In this study, we propose a speaker privacy-preserving representation learning method through the Universal Speech Codec (USC), a computationally efficient encoder-decoder model that disentangles speech into: (i) privacy-preserving semantically rich representations, capturing content and speech paralinguistics, and (ii) residual acoustic and speaker representations that enables high-fidelity reconstruction. Extensive evaluations presented show that USC's semantic representation preserves content, prosody, and sentiment, while removing potentially identifiable speaker attributes. Combining both representations, USC achieves state-of-the-art speech reconstruction. Additionally, we introduce an evaluation methodology for measuring privacy-preserving properties, aligning with perceptual tests. We compare USC against other codecs in the literature and demonstrate its effectiveness on privacy-preserving representation learning, illustrating the trade-offs of speaker anonymization, paralinguistics retention and content preservation in the learned semantic representations. Audio samples are shared in https://www.amazon.science/usc-samples.
title Universal Semantic Disentangled Privacy-preserving Speech Representation Learning
topic Audio and Speech Processing
Machine Learning
url https://arxiv.org/abs/2505.13085