StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866911006563762176 |
|---|---|
| author | Wang, Hui Yang, Yifan Liu, Shujie Li, Jinyu Meng, Lingwei Liu, Yanqing Zhou, Jiaming Sun, Haoqin Lu, Yan Qin, Yong |
| author_facet | Wang, Hui Yang, Yifan Liu, Shujie Li, Jinyu Meng, Lingwei Liu, Yanqing Zhou, Jiaming Sun, Haoqin Lu, Yan Qin, Yong |
| contents | Recent advances in zero-shot text-to-speech (TTS) synthesis have achieved high-quality speech generation for unseen speakers, but most systems remain unsuitable for real-time applications because of their offline design. Current streaming TTS paradigms often rely on multi-stage pipelines and discrete representations, leading to increased computational cost and suboptimal system performance. In this work, we propose StreamMel, a pioneering single-stage streaming TTS framework that models continuous mel-spectrograms. By interleaving text tokens with acoustic frames, StreamMel enables low-latency, autoregressive synthesis while preserving high speaker similarity and naturalness. Experiments on LibriSpeech demonstrate that StreamMel outperforms existing streaming TTS baselines in both quality and latency. It even achieves performance comparable to offline systems while supporting efficient real-time generation, showcasing broad prospects for integration with real-time speech large language models. Audio samples are available at: https://aka.ms/StreamMel. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2506_12570 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling Wang, Hui Yang, Yifan Liu, Shujie Li, Jinyu Meng, Lingwei Liu, Yanqing Zhou, Jiaming Sun, Haoqin Lu, Yan Qin, Yong Sound Computation and Language Audio and Speech Processing Recent advances in zero-shot text-to-speech (TTS) synthesis have achieved high-quality speech generation for unseen speakers, but most systems remain unsuitable for real-time applications because of their offline design. Current streaming TTS paradigms often rely on multi-stage pipelines and discrete representations, leading to increased computational cost and suboptimal system performance. In this work, we propose StreamMel, a pioneering single-stage streaming TTS framework that models continuous mel-spectrograms. By interleaving text tokens with acoustic frames, StreamMel enables low-latency, autoregressive synthesis while preserving high speaker similarity and naturalness. Experiments on LibriSpeech demonstrate that StreamMel outperforms existing streaming TTS baselines in both quality and latency. It even achieves performance comparable to offline systems while supporting efficient real-time generation, showcasing broad prospects for integration with real-time speech large language models. Audio samples are available at: https://aka.ms/StreamMel. |
| title | StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling |
| topic | Sound Computation and Language Audio and Speech Processing |
| url | https://arxiv.org/abs/2506.12570 |