Matcha-TTS: A fast TTS architecture with conditional flow matching

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mehta, Shivam, Tu, Ruibo, Beskow, Jonas, Székely, Éva, Henter, Gustav Eje
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916086524411904
author Mehta, Shivam
Tu, Ruibo
Beskow, Jonas
Székely, Éva
Henter, Gustav Eje
author_facet Mehta, Shivam
Tu, Ruibo
Beskow, Jonas
Székely, Éva
Henter, Gustav Eje
contents We introduce Matcha-TTS, a new encoder-decoder architecture for speedy TTS acoustic modelling, trained using optimal-transport conditional flow matching (OT-CFM). This yields an ODE-based decoder capable of high output quality in fewer synthesis steps than models trained using score matching. Careful design choices additionally ensure each synthesis step is fast to run. The method is probabilistic, non-autoregressive, and learns to speak from scratch without external alignments. Compared to strong pre-trained baseline models, the Matcha-TTS system has the smallest memory footprint, rivals the speed of the fastest models on long utterances, and attains the highest mean opinion score in a listening test. Please see https://shivammehta25.github.io/Matcha-TTS/ for audio examples, code, and pre-trained models.
format Preprint
id arxiv_https___arxiv_org_abs_2309_03199
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Matcha-TTS: A fast TTS architecture with conditional flow matching
Mehta, Shivam
Tu, Ruibo
Beskow, Jonas
Székely, Éva
Henter, Gustav Eje
Audio and Speech Processing
Human-Computer Interaction
Machine Learning
Sound
68T07
I.2.7; I.2.6; H.5.5
We introduce Matcha-TTS, a new encoder-decoder architecture for speedy TTS acoustic modelling, trained using optimal-transport conditional flow matching (OT-CFM). This yields an ODE-based decoder capable of high output quality in fewer synthesis steps than models trained using score matching. Careful design choices additionally ensure each synthesis step is fast to run. The method is probabilistic, non-autoregressive, and learns to speak from scratch without external alignments. Compared to strong pre-trained baseline models, the Matcha-TTS system has the smallest memory footprint, rivals the speed of the fastest models on long utterances, and attains the highest mean opinion score in a listening test. Please see https://shivammehta25.github.io/Matcha-TTS/ for audio examples, code, and pre-trained models.
title Matcha-TTS: A fast TTS architecture with conditional flow matching
topic Audio and Speech Processing
Human-Computer Interaction
Machine Learning
Sound
68T07
I.2.7; I.2.6; H.5.5
url https://arxiv.org/abs/2309.03199