StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Wang, Hui, Yang, Yifan, Liu, Shujie, Li, Jinyu, Meng, Lingwei, Liu, Yanqing, Zhou, Jiaming, Sun, Haoqin, Lu, Yan, Qin, Yong
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866911006563762176
author Wang, Hui
Yang, Yifan
Liu, Shujie
Li, Jinyu
Meng, Lingwei
Liu, Yanqing
Zhou, Jiaming
Sun, Haoqin
Lu, Yan
Qin, Yong
author_facet Wang, Hui
Yang, Yifan
Liu, Shujie
Li, Jinyu
Meng, Lingwei
Liu, Yanqing
Zhou, Jiaming
Sun, Haoqin
Lu, Yan
Qin, Yong
contents Recent advances in zero-shot text-to-speech (TTS) synthesis have achieved high-quality speech generation for unseen speakers, but most systems remain unsuitable for real-time applications because of their offline design. Current streaming TTS paradigms often rely on multi-stage pipelines and discrete representations, leading to increased computational cost and suboptimal system performance. In this work, we propose StreamMel, a pioneering single-stage streaming TTS framework that models continuous mel-spectrograms. By interleaving text tokens with acoustic frames, StreamMel enables low-latency, autoregressive synthesis while preserving high speaker similarity and naturalness. Experiments on LibriSpeech demonstrate that StreamMel outperforms existing streaming TTS baselines in both quality and latency. It even achieves performance comparable to offline systems while supporting efficient real-time generation, showcasing broad prospects for integration with real-time speech large language models. Audio samples are available at: https://aka.ms/StreamMel.
format Preprint
id arxiv_https___arxiv_org_abs_2506_12570
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling
Wang, Hui
Yang, Yifan
Liu, Shujie
Li, Jinyu
Meng, Lingwei
Liu, Yanqing
Zhou, Jiaming
Sun, Haoqin
Lu, Yan
Qin, Yong
Sound
Computation and Language
Audio and Speech Processing
Recent advances in zero-shot text-to-speech (TTS) synthesis have achieved high-quality speech generation for unseen speakers, but most systems remain unsuitable for real-time applications because of their offline design. Current streaming TTS paradigms often rely on multi-stage pipelines and discrete representations, leading to increased computational cost and suboptimal system performance. In this work, we propose StreamMel, a pioneering single-stage streaming TTS framework that models continuous mel-spectrograms. By interleaving text tokens with acoustic frames, StreamMel enables low-latency, autoregressive synthesis while preserving high speaker similarity and naturalness. Experiments on LibriSpeech demonstrate that StreamMel outperforms existing streaming TTS baselines in both quality and latency. It even achieves performance comparable to offline systems while supporting efficient real-time generation, showcasing broad prospects for integration with real-time speech large language models. Audio samples are available at: https://aka.ms/StreamMel.
title StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling
topic Sound
Computation and Language
Audio and Speech Processing
url https://arxiv.org/abs/2506.12570