CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Du, Zhihao, Gao, Changfeng, Wang, Yuxuan, Yu, Fan, Zhao, Tianyu, Wang, Hao, Lv, Xiang, Wang, Hui, Ni, Chongjia, Shi, Xian, An, Keyu, Yang, Guanrou, Li, Yabin, Chen, Yanni, Gao, Zhifu, Chen, Qian, Gu, Yue, Chen, Mengzhe, Chen, Yafeng, Zhang, Shiliang, Wang, Wen, Ye, Jieping
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913860880957440
author Du, Zhihao
Gao, Changfeng
Wang, Yuxuan
Yu, Fan
Zhao, Tianyu
Wang, Hao
Lv, Xiang
Wang, Hui
Ni, Chongjia
Shi, Xian
An, Keyu
Yang, Guanrou
Li, Yabin
Chen, Yanni
Gao, Zhifu
Chen, Qian
Gu, Yue
Chen, Mengzhe
Chen, Yafeng
Zhang, Shiliang
Wang, Wen
Ye, Jieping
author_facet Du, Zhihao
Gao, Changfeng
Wang, Yuxuan
Yu, Fan
Zhao, Tianyu
Wang, Hao
Lv, Xiang
Wang, Hui
Ni, Chongjia
Shi, Xian
An, Keyu
Yang, Guanrou
Li, Yabin
Chen, Yanni
Gao, Zhifu
Chen, Qian
Gu, Yue
Chen, Mengzhe
Chen, Yafeng
Zhang, Shiliang
Wang, Wen
Ye, Jieping
contents In our prior works, we introduced a scalable streaming speech synthesis model, CosyVoice 2, which integrates a large language model (LLM) and a chunk-aware flow matching (FM) model, and achieves low-latency bi-streaming speech synthesis and human-parity quality. Despite these advancements, CosyVoice 2 exhibits limitations in language coverage, domain diversity, data volume, text formats, and post-training techniques. In this paper, we present CosyVoice 3, an improved model designed for zero-shot multilingual speech synthesis in the wild, surpassing its predecessor in content consistency, speaker similarity, and prosody naturalness. Key features of CosyVoice 3 include: 1) A novel speech tokenizer to improve prosody naturalness, developed via supervised multi-task training, including automatic speech recognition, speech emotion recognition, language identification, audio event detection, and speaker analysis. 2) A new differentiable reward model for post-training applicable not only to CosyVoice 3 but also to other LLM-based speech synthesis models. 3) Dataset Size Scaling: Training data is expanded from ten thousand hours to one million hours, encompassing 9 languages and 18 Chinese dialects across various domains and text formats. 4) Model Size Scaling: Model parameters are increased from 0.5 billion to 1.5 billion, resulting in enhanced performance on our multilingual benchmark due to the larger model capacity. These advancements contribute significantly to the progress of speech synthesis in the wild. We encourage readers to listen to the demo at https://funaudiollm.github.io/cosyvoice3.
format Preprint
id arxiv_https___arxiv_org_abs_2505_17589
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training
Du, Zhihao
Gao, Changfeng
Wang, Yuxuan
Yu, Fan
Zhao, Tianyu
Wang, Hao
Lv, Xiang
Wang, Hui
Ni, Chongjia
Shi, Xian
An, Keyu
Yang, Guanrou
Li, Yabin
Chen, Yanni
Gao, Zhifu
Chen, Qian
Gu, Yue
Chen, Mengzhe
Chen, Yafeng
Zhang, Shiliang
Wang, Wen
Ye, Jieping
Sound
Artificial Intelligence
Audio and Speech Processing
In our prior works, we introduced a scalable streaming speech synthesis model, CosyVoice 2, which integrates a large language model (LLM) and a chunk-aware flow matching (FM) model, and achieves low-latency bi-streaming speech synthesis and human-parity quality. Despite these advancements, CosyVoice 2 exhibits limitations in language coverage, domain diversity, data volume, text formats, and post-training techniques. In this paper, we present CosyVoice 3, an improved model designed for zero-shot multilingual speech synthesis in the wild, surpassing its predecessor in content consistency, speaker similarity, and prosody naturalness. Key features of CosyVoice 3 include: 1) A novel speech tokenizer to improve prosody naturalness, developed via supervised multi-task training, including automatic speech recognition, speech emotion recognition, language identification, audio event detection, and speaker analysis. 2) A new differentiable reward model for post-training applicable not only to CosyVoice 3 but also to other LLM-based speech synthesis models. 3) Dataset Size Scaling: Training data is expanded from ten thousand hours to one million hours, encompassing 9 languages and 18 Chinese dialects across various domains and text formats. 4) Model Size Scaling: Model parameters are increased from 0.5 billion to 1.5 billion, resulting in enhanced performance on our multilingual benchmark due to the larger model capacity. These advancements contribute significantly to the progress of speech synthesis in the wild. We encourage readers to listen to the demo at https://funaudiollm.github.io/cosyvoice3.
title CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training
topic Sound
Artificial Intelligence
Audio and Speech Processing
url https://arxiv.org/abs/2505.17589