CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training
Fuente:
arXiv
Saved in:
| Main Authors: | Du, Zhihao, Gao, Changfeng, Wang, Yuxuan, Yu, Fan, Zhao, Tianyu, Wang, Hao, Lv, Xiang, Wang, Hui, Ni, Chongjia, Shi, Xian, An, Keyu, Yang, Guanrou, Li, Yabin, Chen, Yanni, Gao, Zhifu, Chen, Qian, Gu, Yue, Chen, Mengzhe, Chen, Yafeng, Zhang, Shiliang, Wang, Wen, Ye, Jieping |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models
by: Du, Zhihao, et al.
Published: (2024)
by: Du, Zhihao, et al.
Published: (2024)
CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens
by: Du, Zhihao, et al.
Published: (2024)
by: Du, Zhihao, et al.
Published: (2024)
MinMo: A Multimodal Large Language Model for Seamless Voice Interaction
by: Chen, Qian, et al.
Published: (2025)
by: Chen, Qian, et al.
Published: (2025)
Explore the Reinforcement Learning for the LLM based ASR and TTS system
by: Gao, Changfeng, et al.
Published: (2025)
by: Gao, Changfeng, et al.
Published: (2025)
CTC-Assisted LLM-Based Contextual ASR
by: Yang, Guanrou, et al.
Published: (2024)
by: Yang, Guanrou, et al.
Published: (2024)
Enhancing Low-Resource ASR through Versatile TTS: Bridging the Data Gap
by: Yang, Guanrou, et al.
Published: (2024)
by: Yang, Guanrou, et al.
Published: (2024)
MaLa-ASR: Multimedia-Assisted LLM-Based ASR
by: Yang, Guanrou, et al.
Published: (2024)
by: Yang, Guanrou, et al.
Published: (2024)
EmoVoice: LLM-based Emotional Text-To-Speech Model with Freestyle Text Prompting
by: Yang, Guanrou, et al.
Published: (2025)
by: Yang, Guanrou, et al.
Published: (2025)
MELA-TTS: Joint transformer-diffusion model with representation alignment for speech synthesis
by: An, Keyu, et al.
Published: (2025)
by: An, Keyu, et al.
Published: (2025)
Paraformer-v2: An improved non-autoregressive transformer for noise-robust speech recognition
by: An, Keyu, et al.
Published: (2024)
by: An, Keyu, et al.
Published: (2024)
Differentiable Reward Optimization for LLM based TTS system
by: Gao, Changfeng, et al.
Published: (2025)
by: Gao, Changfeng, et al.
Published: (2025)
An Embarrassingly Simple Approach for LLM with Strong ASR Capacity
by: Ma, Ziyang, et al.
Published: (2024)
by: Ma, Ziyang, et al.
Published: (2024)
FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs
by: An, Keyu, et al.
Published: (2024)
by: An, Keyu, et al.
Published: (2024)
DrVoice: Parallel Speech-Text Voice Conversation Model via Dual-Resolution Speech Representations
by: Tan, Chao-Hong, et al.
Published: (2025)
by: Tan, Chao-Hong, et al.
Published: (2025)
CosyEdit: Unlocking End-to-End Speech Editing Capability from Zero-Shot Text-to-Speech Models
by: Chen, Junyang, et al.
Published: (2026)
by: Chen, Junyang, et al.
Published: (2026)
CosyEdit2: Speech-Editing-Oriented Reinforcement Learning Unlocks Better Zero-Shot TTS
by: Chen, Junyang, et al.
Published: (2026)
by: Chen, Junyang, et al.
Published: (2026)
Fun-ASR Technical Report
by: An, Keyu, et al.
Published: (2025)
by: An, Keyu, et al.
Published: (2025)
RRPO: Robust Reward Policy Optimization for LLM-based Emotional TTS
by: Wang, Cong, et al.
Published: (2025)
by: Wang, Cong, et al.
Published: (2025)
DiSTAR: Diffusion over a Scalable Token Autoregressive Representation for Speech Generation
by: Song, Yakun, et al.
Published: (2025)
by: Song, Yakun, et al.
Published: (2025)
TacoLM: GaTed Attention Equipped Codec Language Model are Efficient Zero-Shot Text to Speech Synthesizers
by: Song, Yakun, et al.
Published: (2024)
by: Song, Yakun, et al.
Published: (2024)
Towards Fine-Grained and Multi-Granular Contrastive Language-Speech Pre-training
by: Yang, Yifan, et al.
Published: (2026)
by: Yang, Yifan, et al.
Published: (2026)
Self-Distillation Prototypes Network: Learning Robust Speaker Representations without Supervision
by: Chen, Yafeng, et al.
Published: (2024)
by: Chen, Yafeng, et al.
Published: (2024)
Official Turnover and Embodied Carbon Emissions: Evidence From Industrial Linkages in China's Prefecture‐Level Cities
by: Xuheng Wang, et al.
Published: (2024)
by: Xuheng Wang, et al.
Published: (2024)
Are Transformers in Pre-trained LM A Good ASR Encoder? An Empirical Study
by: An, Keyu, et al.
Published: (2024)
by: An, Keyu, et al.
Published: (2024)
UltraVoice: Scaling Fine-Grained Style-Controlled Speech Conversations for Spoken Dialogue Models
by: Tu, Wenming, et al.
Published: (2025)
by: Tu, Wenming, et al.
Published: (2025)
SyncSpeech: Efficient and Low-Latency Text-to-Speech based on Temporal Masked Transformer
by: Sheng, Zhengyan, et al.
Published: (2025)
by: Sheng, Zhengyan, et al.
Published: (2025)
Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation
by: Liu, Wenrui, et al.
Published: (2025)
by: Liu, Wenrui, et al.
Published: (2025)
On components of the tensor square of a Weyl module
by: Gao, Shiliang, et al.
Published: (2023)
by: Gao, Shiliang, et al.
Published: (2023)
InspireMusic: Integrating Super Resolution and Large Language Model for High-Fidelity Long-Form Music Generation
by: Zhang, Chong, et al.
Published: (2025)
by: Zhang, Chong, et al.
Published: (2025)
Self-Distillation Prototypes Network: Learning Robust Speaker Representations without Supervision
by: Chen, Yafeng, et al.
Published: (2023)
by: Chen, Yafeng, et al.
Published: (2023)
Improving Speaker Diarization using Semantic Information: Joint Pairwise Constraints Propagation
by: Cheng, Luyao, et al.
Published: (2023)
by: Cheng, Luyao, et al.
Published: (2023)
ERes2NetV2: Boosting Short-Duration Speaker Verification Performance with Computational Efficiency
by: Chen, Yafeng, et al.
Published: (2024)
by: Chen, Yafeng, et al.
Published: (2024)
LCB-net: Long-Context Biasing for Audio-Visual Speech Recognition
by: Yu, Fan, et al.
Published: (2024)
by: Yu, Fan, et al.
Published: (2024)
Exploring Efficient Directional and Distance Cues for Regional Speech Separation
by: Jiang, Yiheng, et al.
Published: (2025)
by: Jiang, Yiheng, et al.
Published: (2025)
CosyAudio: Improving Audio Generation with Confidence Scores and Synthetic Captions
by: Zhu, Xinfa, et al.
Published: (2025)
by: Zhu, Xinfa, et al.
Published: (2025)
Interactive ASR: Towards Human-Like Interaction and Semantic Coherence Evaluation for Agentic Speech Recognition
by: Wang, Peng, et al.
Published: (2026)
by: Wang, Peng, et al.
Published: (2026)
A Linear Fractional Transformation Model and Calibration Method for Light Field Camera
by: Chen, Zhong, et al.
Published: (2025)
by: Chen, Zhong, et al.
Published: (2025)
Predicting the low‐cycle fatigue life of Ti‐6Al‐4V alloy using backpropagation neural network optimized by the improved dung beetle algorithm
by: Zihao Gao, et al.
Published: (2024)
by: Zihao Gao, et al.
Published: (2024)
IntrinsicVoice: Empowering LLMs with Intrinsic Real-time Voice Interaction Abilities
by: Zhang, Xin, et al.
Published: (2024)
by: Zhang, Xin, et al.
Published: (2024)
Vevo2: A Unified and Controllable Framework for Speech and Singing Voice Generation
by: Zhang, Xueyao, et al.
Published: (2025)
by: Zhang, Xueyao, et al.
Published: (2025)
Similar Items
-
CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models
by: Du, Zhihao, et al.
Published: (2024) -
CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens
by: Du, Zhihao, et al.
Published: (2024) -
MinMo: A Multimodal Large Language Model for Seamless Voice Interaction
by: Chen, Qian, et al.
Published: (2025) -
Explore the Reinforcement Learning for the LLM based ASR and TTS system
by: Gao, Changfeng, et al.
Published: (2025) -
CTC-Assisted LLM-Based Contextual ASR
by: Yang, Guanrou, et al.
Published: (2024)