NTU-NPU System for Voice Privacy 2024 Challenge
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866914968932188160 |
|---|---|
| author | Kuzmin, Nikita Luong, Hieu-Thi Yao, Jixun Xie, Lei Lee, Kong Aik Chng, Eng Siong |
| author_facet | Kuzmin, Nikita Luong, Hieu-Thi Yao, Jixun Xie, Lei Lee, Kong Aik Chng, Eng Siong |
| contents | In this work, we describe our submissions for the Voice Privacy Challenge 2024. Rather than proposing a novel speech anonymization system, we enhance the provided baselines to meet all required conditions and improve evaluated metrics. Specifically, we implement emotion embedding and experiment with WavLM and ECAPA2 speaker embedders for the B3 baseline. Additionally, we compare different speaker and prosody anonymization techniques. Furthermore, we introduce Mean Reversion F0 for B5, which helps to enhance privacy without a loss in utility. Finally, we explore disentanglement models, namely $β$-VAE and NaturalSpeech3 FACodec. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2410_02371 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | NTU-NPU System for Voice Privacy 2024 Challenge Kuzmin, Nikita Luong, Hieu-Thi Yao, Jixun Xie, Lei Lee, Kong Aik Chng, Eng Siong Audio and Speech Processing Artificial Intelligence In this work, we describe our submissions for the Voice Privacy Challenge 2024. Rather than proposing a novel speech anonymization system, we enhance the provided baselines to meet all required conditions and improve evaluated metrics. Specifically, we implement emotion embedding and experiment with WavLM and ECAPA2 speaker embedders for the B3 baseline. Additionally, we compare different speaker and prosody anonymization techniques. Furthermore, we introduce Mean Reversion F0 for B5, which helps to enhance privacy without a loss in utility. Finally, we explore disentanglement models, namely $β$-VAE and NaturalSpeech3 FACodec. |
| title | NTU-NPU System for Voice Privacy 2024 Challenge |
| topic | Audio and Speech Processing Artificial Intelligence |
| url | https://arxiv.org/abs/2410.02371 |