F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Chen, Yushen, Niu, Zhikang, Ma, Ziyang, Deng, Keqi, Wang, Chunhui, Zhao, Jian, Yu, Kai, Chen, Xie
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866915294300078080
author Chen, Yushen
Niu, Zhikang
Ma, Ziyang
Deng, Keqi
Wang, Chunhui
Zhao, Jian
Yu, Kai
Chen, Xie
author_facet Chen, Yushen
Niu, Zhikang
Ma, Ziyang
Deng, Keqi
Wang, Chunhui
Zhao, Jian
Yu, Kai
Chen, Xie
contents This paper introduces F5-TTS, a fully non-autoregressive text-to-speech system based on flow matching with Diffusion Transformer (DiT). Without requiring complex designs such as duration model, text encoder, and phoneme alignment, the text input is simply padded with filler tokens to the same length as input speech, and then the denoising is performed for speech generation, which was originally proved feasible by E2 TTS. However, the original design of E2 TTS makes it hard to follow due to its slow convergence and low robustness. To address these issues, we first model the input with ConvNeXt to refine the text representation, making it easy to align with the speech. We further propose an inference-time Sway Sampling strategy, which significantly improves our model's performance and efficiency. This sampling strategy for flow step can be easily applied to existing flow matching based models without retraining. Our design allows faster training and achieves an inference RTF of 0.15, which is greatly improved compared to state-of-the-art diffusion-based TTS models. Trained on a public 100K hours multilingual dataset, our F5-TTS exhibits highly natural and expressive zero-shot ability, seamless code-switching capability, and speed control efficiency. We have released all codes and checkpoints to promote community development, at https://SWivid.github.io/F5-TTS/.
format Preprint
id arxiv_https___arxiv_org_abs_2410_06885
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching
Chen, Yushen
Niu, Zhikang
Ma, Ziyang
Deng, Keqi
Wang, Chunhui
Zhao, Jian
Yu, Kai
Chen, Xie
Audio and Speech Processing
Sound
This paper introduces F5-TTS, a fully non-autoregressive text-to-speech system based on flow matching with Diffusion Transformer (DiT). Without requiring complex designs such as duration model, text encoder, and phoneme alignment, the text input is simply padded with filler tokens to the same length as input speech, and then the denoising is performed for speech generation, which was originally proved feasible by E2 TTS. However, the original design of E2 TTS makes it hard to follow due to its slow convergence and low robustness. To address these issues, we first model the input with ConvNeXt to refine the text representation, making it easy to align with the speech. We further propose an inference-time Sway Sampling strategy, which significantly improves our model's performance and efficiency. This sampling strategy for flow step can be easily applied to existing flow matching based models without retraining. Our design allows faster training and achieves an inference RTF of 0.15, which is greatly improved compared to state-of-the-art diffusion-based TTS models. Trained on a public 100K hours multilingual dataset, our F5-TTS exhibits highly natural and expressive zero-shot ability, seamless code-switching capability, and speed control efficiency. We have released all codes and checkpoints to promote community development, at https://SWivid.github.io/F5-TTS/.
title F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2410.06885