A unified front-end framework for English text-to-speech synthesis

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Ying, Zelin, Li, Chen, Dong, Yu, Kong, Qiuqiang, Tian, Qiao, Huo, Yuanyuan, Wang, Yuxuan
Format: Preprint
Veröffentlicht: 2023
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916175812755456
author Ying, Zelin
Li, Chen
Dong, Yu
Kong, Qiuqiang
Tian, Qiao
Huo, Yuanyuan
Wang, Yuxuan
author_facet Ying, Zelin
Li, Chen
Dong, Yu
Kong, Qiuqiang
Tian, Qiao
Huo, Yuanyuan
Wang, Yuxuan
contents The front-end is a critical component of English text-to-speech (TTS) systems, responsible for extracting linguistic features that are essential for a text-to-speech model to synthesize speech, such as prosodies and phonemes. The English TTS front-end typically consists of a text normalization (TN) module, a prosody word prosody phrase (PWPP) module, and a grapheme-to-phoneme (G2P) module. However, current research on the English TTS front-end focuses solely on individual modules, neglecting the interdependence between them and resulting in sub-optimal performance for each module. Therefore, this paper proposes a unified front-end framework that captures the dependencies among the English TTS front-end modules. Extensive experiments have demonstrated that the proposed method achieves state-of-the-art (SOTA) performance in all modules.
format Preprint
id arxiv_https___arxiv_org_abs_2305_10666
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle A unified front-end framework for English text-to-speech synthesis
Ying, Zelin
Li, Chen
Dong, Yu
Kong, Qiuqiang
Tian, Qiao
Huo, Yuanyuan
Wang, Yuxuan
Computation and Language
Artificial Intelligence
Sound
Audio and Speech Processing
The front-end is a critical component of English text-to-speech (TTS) systems, responsible for extracting linguistic features that are essential for a text-to-speech model to synthesize speech, such as prosodies and phonemes. The English TTS front-end typically consists of a text normalization (TN) module, a prosody word prosody phrase (PWPP) module, and a grapheme-to-phoneme (G2P) module. However, current research on the English TTS front-end focuses solely on individual modules, neglecting the interdependence between them and resulting in sub-optimal performance for each module. Therefore, this paper proposes a unified front-end framework that captures the dependencies among the English TTS front-end modules. Extensive experiments have demonstrated that the proposed method achieves state-of-the-art (SOTA) performance in all modules.
title A unified front-end framework for English text-to-speech synthesis
topic Computation and Language
Artificial Intelligence
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2305.10666