FLY-TTS: Fast, Lightweight and High-Quality End-to-End Text-to-Speech Synthesis
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866929404531179520 |
|---|---|
| author | Guo, Yinlin Lv, Yening Dou, Jinqiao Zhang, Yan Wang, Yuehai |
| author_facet | Guo, Yinlin Lv, Yening Dou, Jinqiao Zhang, Yan Wang, Yuehai |
| contents | While recent advances in Text-To-Speech synthesis have yielded remarkable improvements in generating high-quality speech, research on lightweight and fast models is limited. This paper introduces FLY-TTS, a new fast, lightweight and high-quality speech synthesis system based on VITS. Specifically, 1) We replace the decoder with ConvNeXt blocks that generate Fourier spectral coefficients followed by the inverse short-time Fourier transform to synthesize waveforms; 2) To compress the model size, we introduce the grouped parameter-sharing mechanism to the text encoder and flow-based model; 3) We further employ the large pre-trained WavLM model for adversarial training to improve synthesis quality. Experimental results show that our model achieves a real-time factor of 0.0139 on an Intel Core i9 CPU, 8.8x faster than the baseline (0.1221), with a 1.6x parameter compression. Objective and subjective evaluations indicate that FLY-TTS exhibits comparable speech quality to the strong baseline. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2407_00753 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | FLY-TTS: Fast, Lightweight and High-Quality End-to-End Text-to-Speech Synthesis Guo, Yinlin Lv, Yening Dou, Jinqiao Zhang, Yan Wang, Yuehai Audio and Speech Processing Sound While recent advances in Text-To-Speech synthesis have yielded remarkable improvements in generating high-quality speech, research on lightweight and fast models is limited. This paper introduces FLY-TTS, a new fast, lightweight and high-quality speech synthesis system based on VITS. Specifically, 1) We replace the decoder with ConvNeXt blocks that generate Fourier spectral coefficients followed by the inverse short-time Fourier transform to synthesize waveforms; 2) To compress the model size, we introduce the grouped parameter-sharing mechanism to the text encoder and flow-based model; 3) We further employ the large pre-trained WavLM model for adversarial training to improve synthesis quality. Experimental results show that our model achieves a real-time factor of 0.0139 on an Intel Core i9 CPU, 8.8x faster than the baseline (0.1221), with a 1.6x parameter compression. Objective and subjective evaluations indicate that FLY-TTS exhibits comparable speech quality to the strong baseline. |
| title | FLY-TTS: Fast, Lightweight and High-Quality End-to-End Text-to-Speech Synthesis |
| topic | Audio and Speech Processing Sound |
| url | https://arxiv.org/abs/2407.00753 |