ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhu, Han, Kang, Wei, Yao, Zengwei, Guo, Liyong, Kuang, Fangjun, Li, Zhaoqing, Zhuang, Weiji, Lin, Long, Povey, Daniel
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912524448825344
author Zhu, Han
Kang, Wei
Yao, Zengwei
Guo, Liyong
Kuang, Fangjun
Li, Zhaoqing
Zhuang, Weiji
Lin, Long
Povey, Daniel
author_facet Zhu, Han
Kang, Wei
Yao, Zengwei
Guo, Liyong
Kuang, Fangjun
Li, Zhaoqing
Zhuang, Weiji
Lin, Long
Povey, Daniel
contents Existing large-scale zero-shot text-to-speech (TTS) models deliver high speech quality but suffer from slow inference speeds due to massive parameters. To address this issue, this paper introduces ZipVoice, a high-quality flow-matching-based zero-shot TTS model with a compact model size and fast inference speed. Key designs include: 1) a Zipformer-based vector field estimator to maintain adequate modeling capabilities under constrained size; 2) Average upsampling-based initial speech-text alignment and Zipformer-based text encoder to improve speech intelligibility; 3) A flow distillation method to reduce sampling steps and eliminate the inference overhead associated with classifier-free guidance. Experiments on 100k hours multilingual datasets show that ZipVoice matches state-of-the-art models in speech quality, while being 3 times smaller and up to 30 times faster than a DiT-based flow-matching baseline. Codes, model checkpoints and demo samples are publicly available at https://github.com/k2-fsa/ZipVoice.
format Preprint
id arxiv_https___arxiv_org_abs_2506_13053
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching
Zhu, Han
Kang, Wei
Yao, Zengwei
Guo, Liyong
Kuang, Fangjun
Li, Zhaoqing
Zhuang, Weiji
Lin, Long
Povey, Daniel
Audio and Speech Processing
Sound
Existing large-scale zero-shot text-to-speech (TTS) models deliver high speech quality but suffer from slow inference speeds due to massive parameters. To address this issue, this paper introduces ZipVoice, a high-quality flow-matching-based zero-shot TTS model with a compact model size and fast inference speed. Key designs include: 1) a Zipformer-based vector field estimator to maintain adequate modeling capabilities under constrained size; 2) Average upsampling-based initial speech-text alignment and Zipformer-based text encoder to improve speech intelligibility; 3) A flow distillation method to reduce sampling steps and eliminate the inference overhead associated with classifier-free guidance. Experiments on 100k hours multilingual datasets show that ZipVoice matches state-of-the-art models in speech quality, while being 3 times smaller and up to 30 times faster than a DiT-based flow-matching baseline. Codes, model checkpoints and demo samples are publicly available at https://github.com/k2-fsa/ZipVoice.
title ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2506.13053