MUSA: Multi-lingual Speaker Anonymization via Serial Disentanglement

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Yao, Jixun, Wang, Qing, Guo, Pengcheng, Ning, Ziqian, Yang, Yuguang, Pan, Yu, Xie, Lei
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866917723345256448
author Yao, Jixun
Wang, Qing
Guo, Pengcheng
Ning, Ziqian
Yang, Yuguang
Pan, Yu
Xie, Lei
author_facet Yao, Jixun
Wang, Qing
Guo, Pengcheng
Ning, Ziqian
Yang, Yuguang
Pan, Yu
Xie, Lei
contents Speaker anonymization is an effective privacy protection solution designed to conceal the speaker's identity while preserving the linguistic content and para-linguistic information of the original speech. While most prior studies focus solely on a single language, an ideal speaker anonymization system should be capable of handling multiple languages. This paper proposes MUSA, a Multi-lingual Speaker Anonymization approach that employs a serial disentanglement strategy to perform a step-by-step disentanglement from a global time-invariant representation to a temporal time-variant representation. By utilizing semantic distillation and self-supervised speaker distillation, the serial disentanglement strategy can avoid strong inductive biases and exhibit superior generalization performance across different languages. Meanwhile, we propose a straightforward anonymization strategy that employs empty embedding with zero values to simulate the speaker identity concealment process, eliminating the need for conversion to a pseudo-speaker identity and thereby reducing the complexity of speaker anonymization process. Experimental results on VoicePrivacy official datasets and multi-lingual datasets demonstrate that MUSA can effectively protect speaker privacy while preserving linguistic content and para-linguistic information.
format Preprint
id arxiv_https___arxiv_org_abs_2407_11629
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle MUSA: Multi-lingual Speaker Anonymization via Serial Disentanglement
Yao, Jixun
Wang, Qing
Guo, Pengcheng
Ning, Ziqian
Yang, Yuguang
Pan, Yu
Xie, Lei
Audio and Speech Processing
Speaker anonymization is an effective privacy protection solution designed to conceal the speaker's identity while preserving the linguistic content and para-linguistic information of the original speech. While most prior studies focus solely on a single language, an ideal speaker anonymization system should be capable of handling multiple languages. This paper proposes MUSA, a Multi-lingual Speaker Anonymization approach that employs a serial disentanglement strategy to perform a step-by-step disentanglement from a global time-invariant representation to a temporal time-variant representation. By utilizing semantic distillation and self-supervised speaker distillation, the serial disentanglement strategy can avoid strong inductive biases and exhibit superior generalization performance across different languages. Meanwhile, we propose a straightforward anonymization strategy that employs empty embedding with zero values to simulate the speaker identity concealment process, eliminating the need for conversion to a pseudo-speaker identity and thereby reducing the complexity of speaker anonymization process. Experimental results on VoicePrivacy official datasets and multi-lingual datasets demonstrate that MUSA can effectively protect speaker privacy while preserving linguistic content and para-linguistic information.
title MUSA: Multi-lingual Speaker Anonymization via Serial Disentanglement
topic Audio and Speech Processing
url https://arxiv.org/abs/2407.11629