Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | , , , , , , |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866913874252398592 |
|---|---|
| author | Zheng, Qixi Chen, Yushen Niu, Zhikang Ma, Ziyang Wang, Xiaofei Yu, Kai Chen, Xie |
| author_facet | Zheng, Qixi Chen, Yushen Niu, Zhikang Ma, Ziyang Wang, Xiaofei Yu, Kai Chen, Xie |
| contents | Flow-matching-based text-to-speech (TTS) models, such as Voicebox, E2 TTS, and F5-TTS, have attracted significant attention in recent years. These models require multiple sampling steps to reconstruct speech from noise, making inference speed a key challenge. Reducing the number of sampling steps can greatly improve inference efficiency. To this end, we introduce Fast F5-TTS, a training-free approach to accelerate the inference of flow-matching-based TTS models. By inspecting the sampling trajectory of F5-TTS, we identify redundant steps and propose Empirically Pruned Step Sampling (EPSS), a non-uniform time-step sampling strategy that effectively reduces the number of sampling steps. Our approach achieves a 7-step generation with an inference RTF of 0.030 on an NVIDIA RTX 3090 GPU, making it 4 times faster than the original F5-TTS while maintaining comparable performance. Furthermore, EPSS performs well on E2 TTS models, demonstrating its strong generalization ability. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2505_19931 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling Zheng, Qixi Chen, Yushen Niu, Zhikang Ma, Ziyang Wang, Xiaofei Yu, Kai Chen, Xie Audio and Speech Processing Sound Flow-matching-based text-to-speech (TTS) models, such as Voicebox, E2 TTS, and F5-TTS, have attracted significant attention in recent years. These models require multiple sampling steps to reconstruct speech from noise, making inference speed a key challenge. Reducing the number of sampling steps can greatly improve inference efficiency. To this end, we introduce Fast F5-TTS, a training-free approach to accelerate the inference of flow-matching-based TTS models. By inspecting the sampling trajectory of F5-TTS, we identify redundant steps and propose Empirically Pruned Step Sampling (EPSS), a non-uniform time-step sampling strategy that effectively reduces the number of sampling steps. Our approach achieves a 7-step generation with an inference RTF of 0.030 on an NVIDIA RTX 3090 GPU, making it 4 times faster than the original F5-TTS while maintaining comparable performance. Furthermore, EPSS performs well on E2 TTS models, demonstrating its strong generalization ability. |
| title | Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling |
| topic | Audio and Speech Processing Sound |
| url | https://arxiv.org/abs/2505.19931 |