DiFlow-TTS: Compact and Low-Latency Zero-Shot Text-to-Speech with Factorized Discrete Flow Matching

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Nguyen, Ngoc-Son, Tran, Thanh V. T., Huynh-Nguyen, Hieu-Nghia, Hy, Truong-Son, Nguyen, Van
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866911358372544512
author Nguyen, Ngoc-Son
Tran, Thanh V. T.
Huynh-Nguyen, Hieu-Nghia
Hy, Truong-Son
Nguyen, Van
author_facet Nguyen, Ngoc-Son
Tran, Thanh V. T.
Huynh-Nguyen, Hieu-Nghia
Hy, Truong-Son
Nguyen, Van
contents This paper introduces DiFlow-TTS, a novel zero-shot text-to-speech (TTS) system that employs discrete flow matching for generative speech modeling. We position this work as an entry point that may facilitate further advances in this research direction. Through extensive empirical evaluation, we analyze both the strengths and limitations of this approach across key aspects, including naturalness, expressive attributes, speaker identity, and inference latency. To this end, we leverage factorized speech representations and design a deterministic Phoneme-Content Mapper for modeling linguistic content, together with a Factorized Discrete Flow Denoiser that jointly models multiple discrete token streams corresponding to prosody and acoustics to capture expressive speech attributes. Experimental results demonstrate that DiFlow-TTS achieves strong performance across multiple metrics while maintaining a compact model size, up to 11.7 times smaller, and enabling low-latency inference that is up to 34 times faster than recent state-of-the-art baselines. Audio samples are available on our demo page: https://diflow-tts.github.io.
format Preprint
id arxiv_https___arxiv_org_abs_2509_09631
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DiFlow-TTS: Compact and Low-Latency Zero-Shot Text-to-Speech with Factorized Discrete Flow Matching
Nguyen, Ngoc-Son
Tran, Thanh V. T.
Huynh-Nguyen, Hieu-Nghia
Hy, Truong-Son
Nguyen, Van
Sound
Computation and Language
Computer Vision and Pattern Recognition
This paper introduces DiFlow-TTS, a novel zero-shot text-to-speech (TTS) system that employs discrete flow matching for generative speech modeling. We position this work as an entry point that may facilitate further advances in this research direction. Through extensive empirical evaluation, we analyze both the strengths and limitations of this approach across key aspects, including naturalness, expressive attributes, speaker identity, and inference latency. To this end, we leverage factorized speech representations and design a deterministic Phoneme-Content Mapper for modeling linguistic content, together with a Factorized Discrete Flow Denoiser that jointly models multiple discrete token streams corresponding to prosody and acoustics to capture expressive speech attributes. Experimental results demonstrate that DiFlow-TTS achieves strong performance across multiple metrics while maintaining a compact model size, up to 11.7 times smaller, and enabling low-latency inference that is up to 34 times faster than recent state-of-the-art baselines. Audio samples are available on our demo page: https://diflow-tts.github.io.
title DiFlow-TTS: Compact and Low-Latency Zero-Shot Text-to-Speech with Factorized Discrete Flow Matching
topic Sound
Computation and Language
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2509.09631