A unified front-end framework for English text-to-speech synthesis
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866916175812755456 |
|---|---|
| author | Ying, Zelin Li, Chen Dong, Yu Kong, Qiuqiang Tian, Qiao Huo, Yuanyuan Wang, Yuxuan |
| author_facet | Ying, Zelin Li, Chen Dong, Yu Kong, Qiuqiang Tian, Qiao Huo, Yuanyuan Wang, Yuxuan |
| contents | The front-end is a critical component of English text-to-speech (TTS) systems, responsible for extracting linguistic features that are essential for a text-to-speech model to synthesize speech, such as prosodies and phonemes. The English TTS front-end typically consists of a text normalization (TN) module, a prosody word prosody phrase (PWPP) module, and a grapheme-to-phoneme (G2P) module. However, current research on the English TTS front-end focuses solely on individual modules, neglecting the interdependence between them and resulting in sub-optimal performance for each module. Therefore, this paper proposes a unified front-end framework that captures the dependencies among the English TTS front-end modules. Extensive experiments have demonstrated that the proposed method achieves state-of-the-art (SOTA) performance in all modules. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2305_10666 |
| institution | arXiv |
| publishDate | 2023 |
| record_format | arxiv |
| spellingShingle | A unified front-end framework for English text-to-speech synthesis Ying, Zelin Li, Chen Dong, Yu Kong, Qiuqiang Tian, Qiao Huo, Yuanyuan Wang, Yuxuan Computation and Language Artificial Intelligence Sound Audio and Speech Processing The front-end is a critical component of English text-to-speech (TTS) systems, responsible for extracting linguistic features that are essential for a text-to-speech model to synthesize speech, such as prosodies and phonemes. The English TTS front-end typically consists of a text normalization (TN) module, a prosody word prosody phrase (PWPP) module, and a grapheme-to-phoneme (G2P) module. However, current research on the English TTS front-end focuses solely on individual modules, neglecting the interdependence between them and resulting in sub-optimal performance for each module. Therefore, this paper proposes a unified front-end framework that captures the dependencies among the English TTS front-end modules. Extensive experiments have demonstrated that the proposed method achieves state-of-the-art (SOTA) performance in all modules. |
| title | A unified front-end framework for English text-to-speech synthesis |
| topic | Computation and Language Artificial Intelligence Sound Audio and Speech Processing |
| url | https://arxiv.org/abs/2305.10666 |