VoiceFlow: Efficient Text-to-Speech with Rectified Flow Matching

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Guo, Yiwei, Du, Chenpeng, Ma, Ziyang, Chen, Xie, Yu, Kai
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916375996399616
author Guo, Yiwei
Du, Chenpeng
Ma, Ziyang
Chen, Xie
Yu, Kai
author_facet Guo, Yiwei
Du, Chenpeng
Ma, Ziyang
Chen, Xie
Yu, Kai
contents Although diffusion models in text-to-speech have become a popular choice due to their strong generative ability, the intrinsic complexity of sampling from diffusion models harms their efficiency. Alternatively, we propose VoiceFlow, an acoustic model that utilizes a rectified flow matching algorithm to achieve high synthesis quality with a limited number of sampling steps. VoiceFlow formulates the process of generating mel-spectrograms into an ordinary differential equation conditional on text inputs, whose vector field is then estimated. The rectified flow technique then effectively straightens its sampling trajectory for efficient synthesis. Subjective and objective evaluations on both single and multi-speaker corpora showed the superior synthesis quality of VoiceFlow compared to the diffusion counterpart. Ablation studies further verified the validity of the rectified flow technique in VoiceFlow.
format Preprint
id arxiv_https___arxiv_org_abs_2309_05027
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle VoiceFlow: Efficient Text-to-Speech with Rectified Flow Matching
Guo, Yiwei
Du, Chenpeng
Ma, Ziyang
Chen, Xie
Yu, Kai
Audio and Speech Processing
Artificial Intelligence
Human-Computer Interaction
Sound
Although diffusion models in text-to-speech have become a popular choice due to their strong generative ability, the intrinsic complexity of sampling from diffusion models harms their efficiency. Alternatively, we propose VoiceFlow, an acoustic model that utilizes a rectified flow matching algorithm to achieve high synthesis quality with a limited number of sampling steps. VoiceFlow formulates the process of generating mel-spectrograms into an ordinary differential equation conditional on text inputs, whose vector field is then estimated. The rectified flow technique then effectively straightens its sampling trajectory for efficient synthesis. Subjective and objective evaluations on both single and multi-speaker corpora showed the superior synthesis quality of VoiceFlow compared to the diffusion counterpart. Ablation studies further verified the validity of the rectified flow technique in VoiceFlow.
title VoiceFlow: Efficient Text-to-Speech with Rectified Flow Matching
topic Audio and Speech Processing
Artificial Intelligence
Human-Computer Interaction
Sound
url https://arxiv.org/abs/2309.05027