Multi-Speaker Multi-Lingual VQTTS System for LIMMITS 2023 Challenge

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Du, Chenpeng, Guo, Yiwei, Shen, Feiyu, Yu, Kai
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915010653978624
author Du, Chenpeng
Guo, Yiwei
Shen, Feiyu
Yu, Kai
author_facet Du, Chenpeng
Guo, Yiwei
Shen, Feiyu
Yu, Kai
contents In this paper, we describe the systems developed by the SJTU X-LANCE team for LIMMITS 2023 Challenge, and we mainly focus on the winning system on naturalness for track 1. The aim of this challenge is to build a multi-speaker multi-lingual text-to-speech (TTS) system for Marathi, Hindi and Telugu. Each of the languages has a male and a female speaker in the given dataset. In track 1, only 5 hours data from each speaker can be selected to train the TTS model. Our system is based on the recently proposed VQTTS that utilizes VQ acoustic feature rather than mel-spectrogram. We introduce additional speaker embeddings and language embeddings to VQTTS for controlling the speaker and language information. In the cross-lingual evaluations where we need to synthesize speech in a cross-lingual speaker's voice, we provide a native speaker's embedding to the acoustic model and the target speaker's embedding to the vocoder. In the subjective MOS listening test on naturalness, our system achieves 4.77 which ranks first.
format Preprint
id arxiv_https___arxiv_org_abs_2304_13121
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Multi-Speaker Multi-Lingual VQTTS System for LIMMITS 2023 Challenge
Du, Chenpeng
Guo, Yiwei
Shen, Feiyu
Yu, Kai
Sound
Audio and Speech Processing
In this paper, we describe the systems developed by the SJTU X-LANCE team for LIMMITS 2023 Challenge, and we mainly focus on the winning system on naturalness for track 1. The aim of this challenge is to build a multi-speaker multi-lingual text-to-speech (TTS) system for Marathi, Hindi and Telugu. Each of the languages has a male and a female speaker in the given dataset. In track 1, only 5 hours data from each speaker can be selected to train the TTS model. Our system is based on the recently proposed VQTTS that utilizes VQ acoustic feature rather than mel-spectrogram. We introduce additional speaker embeddings and language embeddings to VQTTS for controlling the speaker and language information. In the cross-lingual evaluations where we need to synthesize speech in a cross-lingual speaker's voice, we provide a native speaker's embedding to the acoustic model and the target speaker's embedding to the vocoder. In the subjective MOS listening test on naturalness, our system achieves 4.77 which ranks first.
title Multi-Speaker Multi-Lingual VQTTS System for LIMMITS 2023 Challenge
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2304.13121