Speaker anonymization using neural audio codec language models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Panariello, Michele, Nespoli, Francesco, Todisco, Massimiliano, Evans, Nicholas
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911755583619072
author Panariello, Michele
Nespoli, Francesco
Todisco, Massimiliano
Evans, Nicholas
author_facet Panariello, Michele
Nespoli, Francesco
Todisco, Massimiliano
Evans, Nicholas
contents The vast majority of approaches to speaker anonymization involve the extraction of fundamental frequency estimates, linguistic features and a speaker embedding which is perturbed to obfuscate the speaker identity before an anonymized speech waveform is resynthesized using a vocoder. Recent work has shown that x-vector transformations are difficult to control consistently: other sources of speaker information contained within fundamental frequency and linguistic features are re-entangled upon vocoding, meaning that anonymized speech signals still contain speaker information. We propose an approach based upon neural audio codecs (NACs), which are known to generate high-quality synthetic speech when combined with language models. NACs use quantized codes, which are known to effectively bottleneck speaker-related information: we demonstrate the potential of speaker anonymization systems based on NAC language modeling by applying the evaluation framework of the Voice Privacy Challenge 2022.
format Preprint
id arxiv_https___arxiv_org_abs_2309_14129
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Speaker anonymization using neural audio codec language models
Panariello, Michele
Nespoli, Francesco
Todisco, Massimiliano
Evans, Nicholas
Audio and Speech Processing
Sound
The vast majority of approaches to speaker anonymization involve the extraction of fundamental frequency estimates, linguistic features and a speaker embedding which is perturbed to obfuscate the speaker identity before an anonymized speech waveform is resynthesized using a vocoder. Recent work has shown that x-vector transformations are difficult to control consistently: other sources of speaker information contained within fundamental frequency and linguistic features are re-entangled upon vocoding, meaning that anonymized speech signals still contain speaker information. We propose an approach based upon neural audio codecs (NACs), which are known to generate high-quality synthetic speech when combined with language models. NACs use quantized codes, which are known to effectively bottleneck speaker-related information: we demonstrate the potential of speaker anonymization systems based on NAC language modeling by applying the evaluation framework of the Voice Privacy Challenge 2022.
title Speaker anonymization using neural audio codec language models
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2309.14129