WenetSpeech4TTS: A 12,800-hour Mandarin TTS Corpus for Large Speech Generation Model Benchmark

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Ma, Linhan, Guo, Dake, Song, Kun, Jiang, Yuepeng, Wang, Shuai, Xue, Liumeng, Xu, Weiming, Zhao, Huan, Zhang, Binbin, Xie, Lei
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913397973450752
author Ma, Linhan
Guo, Dake
Song, Kun
Jiang, Yuepeng
Wang, Shuai
Xue, Liumeng
Xu, Weiming
Zhao, Huan
Zhang, Binbin
Xie, Lei
author_facet Ma, Linhan
Guo, Dake
Song, Kun
Jiang, Yuepeng
Wang, Shuai
Xue, Liumeng
Xu, Weiming
Zhao, Huan
Zhang, Binbin
Xie, Lei
contents With the development of large text-to-speech (TTS) models and scale-up of the training data, state-of-the-art TTS systems have achieved impressive performance. In this paper, we present WenetSpeech4TTS, a multi-domain Mandarin corpus derived from the open-sourced WenetSpeech dataset. Tailored for the text-to-speech tasks, we refined WenetSpeech by adjusting segment boundaries, enhancing the audio quality, and eliminating speaker mixing within each segment. Following a more accurate transcription process and quality-based data filtering process, the obtained WenetSpeech4TTS corpus contains $12,800$ hours of paired audio-text data. Furthermore, we have created subsets of varying sizes, categorized by segment quality scores to allow for TTS model training and fine-tuning. VALL-E and NaturalSpeech 2 systems are trained and fine-tuned on these subsets to validate the usability of WenetSpeech4TTS, establishing baselines on benchmark for fair comparison of TTS systems. The corpus and corresponding benchmarks are publicly available on huggingface.
format Preprint
id arxiv_https___arxiv_org_abs_2406_05763
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle WenetSpeech4TTS: A 12,800-hour Mandarin TTS Corpus for Large Speech Generation Model Benchmark
Ma, Linhan
Guo, Dake
Song, Kun
Jiang, Yuepeng
Wang, Shuai
Xue, Liumeng
Xu, Weiming
Zhao, Huan
Zhang, Binbin
Xie, Lei
Audio and Speech Processing
With the development of large text-to-speech (TTS) models and scale-up of the training data, state-of-the-art TTS systems have achieved impressive performance. In this paper, we present WenetSpeech4TTS, a multi-domain Mandarin corpus derived from the open-sourced WenetSpeech dataset. Tailored for the text-to-speech tasks, we refined WenetSpeech by adjusting segment boundaries, enhancing the audio quality, and eliminating speaker mixing within each segment. Following a more accurate transcription process and quality-based data filtering process, the obtained WenetSpeech4TTS corpus contains $12,800$ hours of paired audio-text data. Furthermore, we have created subsets of varying sizes, categorized by segment quality scores to allow for TTS model training and fine-tuning. VALL-E and NaturalSpeech 2 systems are trained and fine-tuned on these subsets to validate the usability of WenetSpeech4TTS, establishing baselines on benchmark for fair comparison of TTS systems. The corpus and corresponding benchmarks are publicly available on huggingface.
title WenetSpeech4TTS: A 12,800-hour Mandarin TTS Corpus for Large Speech Generation Model Benchmark
topic Audio and Speech Processing
url https://arxiv.org/abs/2406.05763