HLTCOE JHU Submission to the Voice Privacy Challenge 2024

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Xinyuan, Henry Li, Cai, Zexin, Garg, Ashi, Duh, Kevin, García-Perera, Leibny Paola, Khudanpur, Sanjeev, Andrews, Nicholas, Wiesner, Matthew
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914950994198528
author Xinyuan, Henry Li
Cai, Zexin
Garg, Ashi
Duh, Kevin
García-Perera, Leibny Paola
Khudanpur, Sanjeev
Andrews, Nicholas
Wiesner, Matthew
author_facet Xinyuan, Henry Li
Cai, Zexin
Garg, Ashi
Duh, Kevin
García-Perera, Leibny Paola
Khudanpur, Sanjeev
Andrews, Nicholas
Wiesner, Matthew
contents We present a number of systems for the Voice Privacy Challenge, including voice conversion based systems such as the kNN-VC method and the WavLM voice Conversion method, and text-to-speech (TTS) based systems including Whisper-VITS. We found that while voice conversion systems better preserve emotional content, they struggle to conceal speaker identity in semi-white-box attack scenarios; conversely, TTS methods perform better at anonymization and worse at emotion preservation. Finally, we propose a random admixture system which seeks to balance out the strengths and weaknesses of the two category of systems, achieving a strong EER of over 40% while maintaining UAR at a respectable 47%.
format Preprint
id arxiv_https___arxiv_org_abs_2409_08913
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle HLTCOE JHU Submission to the Voice Privacy Challenge 2024
Xinyuan, Henry Li
Cai, Zexin
Garg, Ashi
Duh, Kevin
García-Perera, Leibny Paola
Khudanpur, Sanjeev
Andrews, Nicholas
Wiesner, Matthew
Audio and Speech Processing
Machine Learning
We present a number of systems for the Voice Privacy Challenge, including voice conversion based systems such as the kNN-VC method and the WavLM voice Conversion method, and text-to-speech (TTS) based systems including Whisper-VITS. We found that while voice conversion systems better preserve emotional content, they struggle to conceal speaker identity in semi-white-box attack scenarios; conversely, TTS methods perform better at anonymization and worse at emotion preservation. Finally, we propose a random admixture system which seeks to balance out the strengths and weaknesses of the two category of systems, achieving a strong EER of over 40% while maintaining UAR at a respectable 47%.
title HLTCOE JHU Submission to the Voice Privacy Challenge 2024
topic Audio and Speech Processing
Machine Learning
url https://arxiv.org/abs/2409.08913