ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912524448825344 |
|---|---|
| author | Zhu, Han Kang, Wei Yao, Zengwei Guo, Liyong Kuang, Fangjun Li, Zhaoqing Zhuang, Weiji Lin, Long Povey, Daniel |
| author_facet | Zhu, Han Kang, Wei Yao, Zengwei Guo, Liyong Kuang, Fangjun Li, Zhaoqing Zhuang, Weiji Lin, Long Povey, Daniel |
| contents | Existing large-scale zero-shot text-to-speech (TTS) models deliver high speech quality but suffer from slow inference speeds due to massive parameters. To address this issue, this paper introduces ZipVoice, a high-quality flow-matching-based zero-shot TTS model with a compact model size and fast inference speed. Key designs include: 1) a Zipformer-based vector field estimator to maintain adequate modeling capabilities under constrained size; 2) Average upsampling-based initial speech-text alignment and Zipformer-based text encoder to improve speech intelligibility; 3) A flow distillation method to reduce sampling steps and eliminate the inference overhead associated with classifier-free guidance. Experiments on 100k hours multilingual datasets show that ZipVoice matches state-of-the-art models in speech quality, while being 3 times smaller and up to 30 times faster than a DiT-based flow-matching baseline. Codes, model checkpoints and demo samples are publicly available at https://github.com/k2-fsa/ZipVoice. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2506_13053 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Zhu, Han Kang, Wei Yao, Zengwei Guo, Liyong Kuang, Fangjun Li, Zhaoqing Zhuang, Weiji Lin, Long Povey, Daniel Audio and Speech Processing Sound Existing large-scale zero-shot text-to-speech (TTS) models deliver high speech quality but suffer from slow inference speeds due to massive parameters. To address this issue, this paper introduces ZipVoice, a high-quality flow-matching-based zero-shot TTS model with a compact model size and fast inference speed. Key designs include: 1) a Zipformer-based vector field estimator to maintain adequate modeling capabilities under constrained size; 2) Average upsampling-based initial speech-text alignment and Zipformer-based text encoder to improve speech intelligibility; 3) A flow distillation method to reduce sampling steps and eliminate the inference overhead associated with classifier-free guidance. Experiments on 100k hours multilingual datasets show that ZipVoice matches state-of-the-art models in speech quality, while being 3 times smaller and up to 30 times faster than a DiT-based flow-matching baseline. Codes, model checkpoints and demo samples are publicly available at https://github.com/k2-fsa/ZipVoice. |
| title | ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching |
| topic | Audio and Speech Processing Sound |
| url | https://arxiv.org/abs/2506.13053 |