Multi-modal Adversarial Training for Zero-Shot Voice Cloning
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866929478021677056 |
|---|---|
| author | Janiczek, John Chong, Dading Dai, Dongyang Faria, Arlo Wang, Chao Wang, Tao Liu, Yuzong |
| author_facet | Janiczek, John Chong, Dading Dai, Dongyang Faria, Arlo Wang, Chao Wang, Tao Liu, Yuzong |
| contents | A text-to-speech (TTS) model trained to reconstruct speech given text tends towards predictions that are close to the average characteristics of a dataset, failing to model the variations that make human speech sound natural. This problem is magnified for zero-shot voice cloning, a task that requires training data with high variance in speaking styles. We build off of recent works which have used Generative Advsarial Networks (GAN) by proposing a Transformer encoder-decoder architecture to conditionally discriminates between real and generated speech features. The discriminator is used in a training pipeline that improves both the acoustic and prosodic features of a TTS model. We introduce our novel adversarial training technique by applying it to a FastSpeech2 acoustic model and training on Libriheavy, a large multi-speaker dataset, for the task of zero-shot voice cloning. Our model achieves improvements over the baseline in terms of speech quality and speaker similarity. Audio examples from our system are available online. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2408_15916 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Multi-modal Adversarial Training for Zero-Shot Voice Cloning Janiczek, John Chong, Dading Dai, Dongyang Faria, Arlo Wang, Chao Wang, Tao Liu, Yuzong Audio and Speech Processing Machine Learning Sound A text-to-speech (TTS) model trained to reconstruct speech given text tends towards predictions that are close to the average characteristics of a dataset, failing to model the variations that make human speech sound natural. This problem is magnified for zero-shot voice cloning, a task that requires training data with high variance in speaking styles. We build off of recent works which have used Generative Advsarial Networks (GAN) by proposing a Transformer encoder-decoder architecture to conditionally discriminates between real and generated speech features. The discriminator is used in a training pipeline that improves both the acoustic and prosodic features of a TTS model. We introduce our novel adversarial training technique by applying it to a FastSpeech2 acoustic model and training on Libriheavy, a large multi-speaker dataset, for the task of zero-shot voice cloning. Our model achieves improvements over the baseline in terms of speech quality and speaker similarity. Audio examples from our system are available online. |
| title | Multi-modal Adversarial Training for Zero-Shot Voice Cloning |
| topic | Audio and Speech Processing Machine Learning Sound |
| url | https://arxiv.org/abs/2408.15916 |