CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Du, Zhihao, Chen, Qian, Zhang, Shiliang, Hu, Kai, Lu, Heng, Yang, Yexin, Hu, Hangrui, Zheng, Siqi, Gu, Yue, Ma, Ziyang, Gao, Zhifu, Yan, Zhijie |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models
par: Du, Zhihao, et autres
Publié: (2024)
par: Du, Zhihao, et autres
Publié: (2024)
CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training
par: Du, Zhihao, et autres
Publié: (2025)
par: Du, Zhihao, et autres
Publié: (2025)
Paraformer-v2: An improved non-autoregressive transformer for noise-robust speech recognition
par: An, Keyu, et autres
Publié: (2024)
par: An, Keyu, et autres
Publié: (2024)
FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs
par: An, Keyu, et autres
Publié: (2024)
par: An, Keyu, et autres
Publié: (2024)
FreeSVC: Towards Zero-shot Multilingual Singing Voice Conversion
par: Ferreira, Alef Iury Siqueira, et autres
Publié: (2025)
par: Ferreira, Alef Iury Siqueira, et autres
Publié: (2025)
Selective Classifier-free Guidance for Zero-shot Text-to-speech
par: Zheng, John, et autres
Publié: (2025)
par: Zheng, John, et autres
Publié: (2025)
TCSinger 2: Customizable Multilingual Zero-shot Singing Voice Synthesis
par: Zhang, Yu, et autres
Publié: (2025)
par: Zhang, Yu, et autres
Publié: (2025)
CosyEdit: Unlocking End-to-End Speech Editing Capability from Zero-Shot Text-to-Speech Models
par: Chen, Junyang, et autres
Publié: (2026)
par: Chen, Junyang, et autres
Publié: (2026)
Enhancing Low-Resource ASR through Versatile TTS: Bridging the Data Gap
par: Yang, Guanrou, et autres
Publié: (2024)
par: Yang, Guanrou, et autres
Publié: (2024)
CTC-Assisted LLM-Based Contextual ASR
par: Yang, Guanrou, et autres
Publié: (2024)
par: Yang, Guanrou, et autres
Publié: (2024)
MaLa-ASR: Multimedia-Assisted LLM-Based ASR
par: Yang, Guanrou, et autres
Publié: (2024)
par: Yang, Guanrou, et autres
Publié: (2024)
StyleFusion TTS: Multimodal Style-control and Enhanced Feature Fusion for Zero-shot Text-to-speech Synthesis
par: Chen, Zhiyong, et autres
Publié: (2024)
par: Chen, Zhiyong, et autres
Publié: (2024)
StreamVoice+: Evolving into End-to-end Streaming Zero-shot Voice Conversion
par: Wang, Zhichao, et autres
Publié: (2024)
par: Wang, Zhichao, et autres
Publié: (2024)
Zero-shot Cross-lingual Voice Transfer for TTS
par: Biadsy, Fadi, et autres
Publié: (2024)
par: Biadsy, Fadi, et autres
Publié: (2024)
LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT
par: Du, Zhihao, et autres
Publié: (2023)
par: Du, Zhihao, et autres
Publié: (2023)
EmoVoice: LLM-based Emotional Text-To-Speech Model with Freestyle Text Prompting
par: Yang, Guanrou, et autres
Publié: (2025)
par: Yang, Guanrou, et autres
Publié: (2025)
Erasing Your Voice Before It's Heard: Training-free Speaker Unlearning for Zero-shot Text-to-Speech
par: Lee, Myungjin, et autres
Publié: (2026)
par: Lee, Myungjin, et autres
Publié: (2026)
BFA: Real-time Multilingual Text-to-speech Forced Alignment
par: Rehman, Abdul, et autres
Publié: (2025)
par: Rehman, Abdul, et autres
Publié: (2025)
Zero-shot Voice Conversion with Diffusion Transformers
par: Liu, Songting
Publié: (2024)
par: Liu, Songting
Publié: (2024)
MELA-TTS: Joint transformer-diffusion model with representation alignment for speech synthesis
par: An, Keyu, et autres
Publié: (2025)
par: An, Keyu, et autres
Publié: (2025)
VoxMorph: Scalable Zero-shot Voice Identity Morphing via Disentangled Embeddings
par: Krishnamurthy, Bharath, et autres
Publié: (2026)
par: Krishnamurthy, Bharath, et autres
Publié: (2026)
Enhancing Zero-shot Text-to-Speech Synthesis with Human Feedback
par: Chen, Chen, et autres
Publié: (2024)
par: Chen, Chen, et autres
Publié: (2024)
An Embarrassingly Simple Approach for LLM with Strong ASR Capacity
par: Ma, Ziyang, et autres
Publié: (2024)
par: Ma, Ziyang, et autres
Publié: (2024)
OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion
par: Wang, Zhichao, et autres
Publié: (2026)
par: Wang, Zhichao, et autres
Publié: (2026)
Improving Generalization for AI-Synthesized Voice Detection
par: Ren, Hainan, et autres
Publié: (2024)
par: Ren, Hainan, et autres
Publié: (2024)
SLM-S2ST: A multimodal language model for direct speech-to-speech translation
par: Hu, Yuxuan, et autres
Publié: (2025)
par: Hu, Yuxuan, et autres
Publié: (2025)
Are Transformers in Pre-trained LM A Good ASR Encoder? An Empirical Study
par: An, Keyu, et autres
Publié: (2024)
par: An, Keyu, et autres
Publié: (2024)
Long-Context Speech Synthesis with Context-Aware Memory
par: Li, Zhipeng, et autres
Publié: (2025)
par: Li, Zhipeng, et autres
Publié: (2025)
CosyAudio: Improving Audio Generation with Confidence Scores and Synthetic Captions
par: Zhu, Xinfa, et autres
Publié: (2025)
par: Zhu, Xinfa, et autres
Publié: (2025)
Enhancing the Robustness of Contextual ASR to Varying Biasing Information Volumes Through Purified Semantic Correlation Joint Modeling
par: Gu, Yue, et autres
Publié: (2025)
par: Gu, Yue, et autres
Publié: (2025)
Multi-level Temporal-channel Speaker Retrieval for Zero-shot Voice Conversion
par: Wang, Zhichao, et autres
Publié: (2023)
par: Wang, Zhichao, et autres
Publié: (2023)
X-VC: Zero-shot Streaming Voice Conversion in Codec Space
par: Zheng, Qixi, et autres
Publié: (2026)
par: Zheng, Qixi, et autres
Publié: (2026)
SSR-Speech: Towards Stable, Safe and Robust Zero-shot Text-based Speech Editing and Synthesis
par: Wang, Helin, et autres
Publié: (2024)
par: Wang, Helin, et autres
Publié: (2024)
EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion
par: Joglekar, Advait, et autres
Publié: (2025)
par: Joglekar, Advait, et autres
Publié: (2025)
Towards Weakly Supervised Text-to-Audio Grounding
par: Xu, Xuenan, et autres
Publié: (2024)
par: Xu, Xuenan, et autres
Publié: (2024)
GenVC: Self-Supervised Zero-Shot Voice Conversion
par: Cai, Zexin, et autres
Publié: (2025)
par: Cai, Zexin, et autres
Publié: (2025)
HAM-TTS: Hierarchical Acoustic Modeling for Token-Based Zero-Shot Text-to-Speech with Model and Data Scaling
par: Wang, Chunhui, et autres
Publié: (2024)
par: Wang, Chunhui, et autres
Publié: (2024)
Self-Distillation Prototypes Network: Learning Robust Speaker Representations without Supervision
par: Chen, Yafeng, et autres
Publié: (2024)
par: Chen, Yafeng, et autres
Publié: (2024)
ZSVC: Zero-shot Style Voice Conversion with Disentangled Latent Diffusion Models and Adversarial Training
par: Zhu, Xinfa, et autres
Publié: (2025)
par: Zhu, Xinfa, et autres
Publié: (2025)
CoDiff-VC: A Codec-Assisted Diffusion Model for Zero-shot Voice Conversion
par: Li, Yuke, et autres
Publié: (2024)
par: Li, Yuke, et autres
Publié: (2024)
Documents similaires
-
CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models
par: Du, Zhihao, et autres
Publié: (2024) -
CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training
par: Du, Zhihao, et autres
Publié: (2025) -
Paraformer-v2: An improved non-autoregressive transformer for noise-robust speech recognition
par: An, Keyu, et autres
Publié: (2024) -
FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs
par: An, Keyu, et autres
Publié: (2024) -
FreeSVC: Towards Zero-shot Multilingual Singing Voice Conversion
par: Ferreira, Alef Iury Siqueira, et autres
Publié: (2025)