ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching
Fuente:
arXiv
Guardado en:
| Autores principales: | Zhu, Han, Kang, Wei, Yao, Zengwei, Guo, Liyong, Kuang, Fangjun, Li, Zhaoqing, Zhuang, Weiji, Lin, Long, Povey, Daniel |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
ZipVoice-Dialog: Non-Autoregressive Spoken Dialogue Generation with Flow Matching
por: Zhu, Han, et al.
Publicado: (2025)
por: Zhu, Han, et al.
Publicado: (2025)
OmniVoice: Towards Omnilingual Zero-Shot Text-to-Speech with Diffusion Language Models
por: Zhu, Han, et al.
Publicado: (2026)
por: Zhu, Han, et al.
Publicado: (2026)
Libriheavy: a 50,000 hours ASR corpus with punctuation casing and context
por: Kang, Wei, et al.
Publicado: (2023)
por: Kang, Wei, et al.
Publicado: (2023)
Flow2GAN: Hybrid Flow Matching and GAN with Multi-Resolution Network for Few-step High-Fidelity Audio Generation
por: Yao, Zengwei, et al.
Publicado: (2025)
por: Yao, Zengwei, et al.
Publicado: (2025)
CR-CTC: Consistency regularization on CTC for improved speech recognition
por: Yao, Zengwei, et al.
Publicado: (2024)
por: Yao, Zengwei, et al.
Publicado: (2024)
PromptASR for contextualized ASR with controllable style
por: Yang, Xiaoyu, et al.
Publicado: (2023)
por: Yang, Xiaoyu, et al.
Publicado: (2023)
Zipformer: A faster and better encoder for automatic speech recognition
por: Yao, Zengwei, et al.
Publicado: (2023)
por: Yao, Zengwei, et al.
Publicado: (2023)
LibriheavyMix: A 20,000-Hour Dataset for Single-Channel Reverberant Multi-Talker Speech Separation, ASR and Speaker Diarization
por: Jin, Zengrui, et al.
Publicado: (2024)
por: Jin, Zengrui, et al.
Publicado: (2024)
SF-Speech: Straightened Flow for Zero-Shot Voice Clone
por: Li, Xuyuan, et al.
Publicado: (2024)
por: Li, Xuyuan, et al.
Publicado: (2024)
StableVC: Style Controllable Zero-Shot Voice Conversion with Conditional Flow Matching
por: Yao, Jixun, et al.
Publicado: (2024)
por: Yao, Jixun, et al.
Publicado: (2024)
VoiceRestore: Flow-Matching Transformers for Speech Recording Quality Restoration
por: Kirdey, Stanislav
Publicado: (2025)
por: Kirdey, Stanislav
Publicado: (2025)
Rhythm Controllable and Efficient Zero-Shot Voice Conversion via Shortcut Flow Matching
por: Zuo, Jialong, et al.
Publicado: (2025)
por: Zuo, Jialong, et al.
Publicado: (2025)
MobileSpeech: A Fast and High-Fidelity Framework for Mobile Zero-Shot Text-to-Speech
por: Ji, Shengpeng, et al.
Publicado: (2024)
por: Ji, Shengpeng, et al.
Publicado: (2024)
Zero-Shot Voice Conversion via Content-Aware Timbre Ensemble and Conditional Flow Matching
por: Pan, Yu, et al.
Publicado: (2024)
por: Pan, Yu, et al.
Publicado: (2024)
MeanVC: Lightweight and Streaming Zero-Shot Voice Conversion via Mean Flows
por: Ma, Guobin, et al.
Publicado: (2025)
por: Ma, Guobin, et al.
Publicado: (2025)
Zero-Shot Text-to-Speech from Continuous Text Streams
por: Dang, Trung, et al.
Publicado: (2024)
por: Dang, Trung, et al.
Publicado: (2024)
End-to-End Zero-Shot Voice Conversion with Location-Variable Convolutions
por: Kang, Wonjune, et al.
Publicado: (2022)
por: Kang, Wonjune, et al.
Publicado: (2022)
Everyone-Can-Sing: Zero-Shot Singing Voice Synthesis and Conversion with Speech Reference
por: Dai, Shuqi, et al.
Publicado: (2025)
por: Dai, Shuqi, et al.
Publicado: (2025)
Do Not Mimic My Voice: Speaker Identity Unlearning for Zero-Shot Text-to-Speech
por: Kim, Taesoo, et al.
Publicado: (2025)
por: Kim, Taesoo, et al.
Publicado: (2025)
Improvement Speaker Similarity for Zero-Shot Any-to-Any Voice Conversion of Whispered and Regular Speech
por: Avdeeva, Anastasia, et al.
Publicado: (2024)
por: Avdeeva, Anastasia, et al.
Publicado: (2024)
Zero Shot Text to Speech Augmentation for Automatic Speech Recognition on Low-Resource Accented Speech Corpora
por: Nespoli, Francesco, et al.
Publicado: (2024)
por: Nespoli, Francesco, et al.
Publicado: (2024)
Advanced Zero-Shot Text-to-Speech for Background Removal and Preservation with Controllable Masked Speech Prediction
por: Zhang, Leying, et al.
Publicado: (2025)
por: Zhang, Leying, et al.
Publicado: (2025)
MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder
por: Zhang, Bowen, et al.
Publicado: (2025)
por: Zhang, Bowen, et al.
Publicado: (2025)
VoicePrompter: Robust Zero-Shot Voice Conversion with Voice Prompt and Conditional Flow Matching
por: Choi, Ha-Yeong, et al.
Publicado: (2025)
por: Choi, Ha-Yeong, et al.
Publicado: (2025)
Zero-Shot Text-to-Speech for Vietnamese
por: Vu, Thi, et al.
Publicado: (2025)
por: Vu, Thi, et al.
Publicado: (2025)
VoiceCraft: Zero-Shot Speech Editing and Text-to-Speech in the Wild
por: Peng, Puyuan, et al.
Publicado: (2024)
por: Peng, Puyuan, et al.
Publicado: (2024)
Improving Language Model-Based Zero-Shot Text-to-Speech Synthesis with Multi-Scale Acoustic Prompts
por: Lei, Shun, et al.
Publicado: (2023)
por: Lei, Shun, et al.
Publicado: (2023)
Word-Level Emotional Expression Control in Zero-Shot Text-to-Speech Synthesis
por: Wang, Tianrui, et al.
Publicado: (2025)
por: Wang, Tianrui, et al.
Publicado: (2025)
StyleTTS-ZS: Efficient High-Quality Zero-Shot Text-to-Speech Synthesis with Distilled Time-Varying Style Diffusion
por: Li, Yinghao Aaron, et al.
Publicado: (2024)
por: Li, Yinghao Aaron, et al.
Publicado: (2024)
VoiceShop: A Unified Speech-to-Speech Framework for Identity-Preserving Zero-Shot Voice Editing
por: Anastassiou, Philip, et al.
Publicado: (2024)
por: Anastassiou, Philip, et al.
Publicado: (2024)
MPE-TTS: Customized Emotion Zero-Shot Text-To-Speech Using Multi-Modal Prompt
por: Wu, Zhichao, et al.
Publicado: (2025)
por: Wu, Zhichao, et al.
Publicado: (2025)
Making Flow-Matching-Based Zero-Shot Text-to-Speech Laugh as You Like
por: Kanda, Naoyuki, et al.
Publicado: (2024)
por: Kanda, Naoyuki, et al.
Publicado: (2024)
k2SSL: A Faster and Better Framework for Self-Supervised Speech Representation Learning
por: Yang, Yifan, et al.
Publicado: (2024)
por: Yang, Yifan, et al.
Publicado: (2024)
CosyEdit: Unlocking End-to-End Speech Editing Capability from Zero-Shot Text-to-Speech Models
por: Chen, Junyang, et al.
Publicado: (2026)
por: Chen, Junyang, et al.
Publicado: (2026)
Erasing Your Voice Before It's Heard: Training-free Speaker Unlearning for Zero-shot Text-to-Speech
por: Lee, Myungjin, et al.
Publicado: (2026)
por: Lee, Myungjin, et al.
Publicado: (2026)
Zero-Shot Duet Singing Voices Separation with Diffusion Models
por: Yu, Chin-Yun, et al.
Publicado: (2023)
por: Yu, Chin-Yun, et al.
Publicado: (2023)
VoiceFlow: Efficient Text-to-Speech with Rectified Flow Matching
por: Guo, Yiwei, et al.
Publicado: (2023)
por: Guo, Yiwei, et al.
Publicado: (2023)
CLaM-TTS: Improving Neural Codec Language Model for Zero-Shot Text-to-Speech
por: Kim, Jaehyeon, et al.
Publicado: (2024)
por: Kim, Jaehyeon, et al.
Publicado: (2024)
Parallel GPT: Harmonizing the Independence and Interdependence of Acoustic and Semantic Information for Zero-Shot Text-to-Speech
por: Xing, Jingyuan, et al.
Publicado: (2025)
por: Xing, Jingyuan, et al.
Publicado: (2025)
Speech Synthesis along Perceptual Voice Quality Dimensions
por: Rautenberg, Frederik, et al.
Publicado: (2025)
por: Rautenberg, Frederik, et al.
Publicado: (2025)
Ejemplares similares
-
ZipVoice-Dialog: Non-Autoregressive Spoken Dialogue Generation with Flow Matching
por: Zhu, Han, et al.
Publicado: (2025) -
OmniVoice: Towards Omnilingual Zero-Shot Text-to-Speech with Diffusion Language Models
por: Zhu, Han, et al.
Publicado: (2026) -
Libriheavy: a 50,000 hours ASR corpus with punctuation casing and context
por: Kang, Wei, et al.
Publicado: (2023) -
Flow2GAN: Hybrid Flow Matching and GAN with Multi-Resolution Network for Few-step High-Fidelity Audio Generation
por: Yao, Zengwei, et al.
Publicado: (2025) -
CR-CTC: Consistency regularization on CTC for improved speech recognition
por: Yao, Zengwei, et al.
Publicado: (2024)