Multi-modal Adversarial Training for Zero-Shot Voice Cloning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Janiczek, John, Chong, Dading, Dai, Dongyang, Faria, Arlo, Wang, Chao, Wang, Tao, Liu, Yuzong
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929478021677056
author Janiczek, John
Chong, Dading
Dai, Dongyang
Faria, Arlo
Wang, Chao
Wang, Tao
Liu, Yuzong
author_facet Janiczek, John
Chong, Dading
Dai, Dongyang
Faria, Arlo
Wang, Chao
Wang, Tao
Liu, Yuzong
contents A text-to-speech (TTS) model trained to reconstruct speech given text tends towards predictions that are close to the average characteristics of a dataset, failing to model the variations that make human speech sound natural. This problem is magnified for zero-shot voice cloning, a task that requires training data with high variance in speaking styles. We build off of recent works which have used Generative Advsarial Networks (GAN) by proposing a Transformer encoder-decoder architecture to conditionally discriminates between real and generated speech features. The discriminator is used in a training pipeline that improves both the acoustic and prosodic features of a TTS model. We introduce our novel adversarial training technique by applying it to a FastSpeech2 acoustic model and training on Libriheavy, a large multi-speaker dataset, for the task of zero-shot voice cloning. Our model achieves improvements over the baseline in terms of speech quality and speaker similarity. Audio examples from our system are available online.
format Preprint
id arxiv_https___arxiv_org_abs_2408_15916
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Multi-modal Adversarial Training for Zero-Shot Voice Cloning
Janiczek, John
Chong, Dading
Dai, Dongyang
Faria, Arlo
Wang, Chao
Wang, Tao
Liu, Yuzong
Audio and Speech Processing
Machine Learning
Sound
A text-to-speech (TTS) model trained to reconstruct speech given text tends towards predictions that are close to the average characteristics of a dataset, failing to model the variations that make human speech sound natural. This problem is magnified for zero-shot voice cloning, a task that requires training data with high variance in speaking styles. We build off of recent works which have used Generative Advsarial Networks (GAN) by proposing a Transformer encoder-decoder architecture to conditionally discriminates between real and generated speech features. The discriminator is used in a training pipeline that improves both the acoustic and prosodic features of a TTS model. We introduce our novel adversarial training technique by applying it to a FastSpeech2 acoustic model and training on Libriheavy, a large multi-speaker dataset, for the task of zero-shot voice cloning. Our model achieves improvements over the baseline in terms of speech quality and speaker similarity. Audio examples from our system are available online.
title Multi-modal Adversarial Training for Zero-Shot Voice Cloning
topic Audio and Speech Processing
Machine Learning
Sound
url https://arxiv.org/abs/2408.15916