FlashSpeech: Efficient Zero-Shot Speech Synthesis
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ye, Zhen, Ju, Zeqian, Liu, Haohe, Tan, Xu, Chen, Jianyi, Lu, Yiwen, Sun, Peiwen, Pan, Jiahao, Bian, Weizhen, He, Shulin, Xue, Wei, Liu, Qifeng, Guo, Yike |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models
von: Ju, Zeqian, et al.
Veröffentlicht: (2024)
von: Ju, Zeqian, et al.
Veröffentlicht: (2024)
Mega-TTS 2: Boosting Prompting Mechanisms for Zero-Shot Speech Synthesis
von: Jiang, Ziyue, et al.
Veröffentlicht: (2023)
von: Jiang, Ziyue, et al.
Veröffentlicht: (2023)
EmoSpeech: A Corpus of Emotionally Rich and Contextually Detailed Speech Annotations
von: Bian, Weizhen, et al.
Veröffentlicht: (2024)
von: Bian, Weizhen, et al.
Veröffentlicht: (2024)
Towards Expressive Zero-Shot Speech Synthesis with Hierarchical Prosody Modeling
von: Jiang, Yuepeng, et al.
Veröffentlicht: (2024)
von: Jiang, Yuepeng, et al.
Veröffentlicht: (2024)
Zero-Shot Audio Captioning Using Soft and Hard Prompts
von: Zhang, Yiming, et al.
Veröffentlicht: (2024)
von: Zhang, Yiming, et al.
Veröffentlicht: (2024)
Word-Level Emotional Expression Control in Zero-Shot Text-to-Speech Synthesis
von: Wang, Tianrui, et al.
Veröffentlicht: (2025)
von: Wang, Tianrui, et al.
Veröffentlicht: (2025)
Zero-Shot Mono-to-Binaural Speech Synthesis
von: Levkovitch, Alon, et al.
Veröffentlicht: (2024)
von: Levkovitch, Alon, et al.
Veröffentlicht: (2024)
Zero Shot Text to Speech Augmentation for Automatic Speech Recognition on Low-Resource Accented Speech Corpora
von: Nespoli, Francesco, et al.
Veröffentlicht: (2024)
von: Nespoli, Francesco, et al.
Veröffentlicht: (2024)
FastSAG: Towards Fast Non-Autoregressive Singing Accompaniment Generation
von: Chen, Jianyi, et al.
Veröffentlicht: (2024)
von: Chen, Jianyi, et al.
Veröffentlicht: (2024)
Everyone-Can-Sing: Zero-Shot Singing Voice Synthesis and Conversion with Speech Reference
von: Dai, Shuqi, et al.
Veröffentlicht: (2025)
von: Dai, Shuqi, et al.
Veröffentlicht: (2025)
CoMoSVC: Consistency Model-based Singing Voice Conversion
von: Lu, Yiwen, et al.
Veröffentlicht: (2024)
von: Lu, Yiwen, et al.
Veröffentlicht: (2024)
Advanced Zero-Shot Text-to-Speech for Background Removal and Preservation with Controllable Masked Speech Prediction
von: Zhang, Leying, et al.
Veröffentlicht: (2025)
von: Zhang, Leying, et al.
Veröffentlicht: (2025)
MobileSpeech: A Fast and High-Fidelity Framework for Mobile Zero-Shot Text-to-Speech
von: Ji, Shengpeng, et al.
Veröffentlicht: (2024)
von: Ji, Shengpeng, et al.
Veröffentlicht: (2024)
MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder
von: Zhang, Bowen, et al.
Veröffentlicht: (2025)
von: Zhang, Bowen, et al.
Veröffentlicht: (2025)
Improving Robustness of Diffusion-Based Zero-Shot Speech Synthesis via Stable Formant Generation
von: Han, Changjin, et al.
Veröffentlicht: (2024)
von: Han, Changjin, et al.
Veröffentlicht: (2024)
SF-Speech: Straightened Flow for Zero-Shot Voice Clone
von: Li, Xuyuan, et al.
Veröffentlicht: (2024)
von: Li, Xuyuan, et al.
Veröffentlicht: (2024)
Zero-Shot Text-to-Speech from Continuous Text Streams
von: Dang, Trung, et al.
Veröffentlicht: (2024)
von: Dang, Trung, et al.
Veröffentlicht: (2024)
Codec Does Matter: Exploring the Semantic Shortcoming of Codec for Audio Language Model
von: Ye, Zhen, et al.
Veröffentlicht: (2024)
von: Ye, Zhen, et al.
Veröffentlicht: (2024)
Improving Language Model-Based Zero-Shot Text-to-Speech Synthesis with Multi-Scale Acoustic Prompts
von: Lei, Shun, et al.
Veröffentlicht: (2023)
von: Lei, Shun, et al.
Veröffentlicht: (2023)
ToneUnit: A Speech Discretization Approach for Tonal Language Speech Synthesis
von: Tao, Dehua, et al.
Veröffentlicht: (2024)
von: Tao, Dehua, et al.
Veröffentlicht: (2024)
Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
von: Wang, Xinsheng, et al.
Veröffentlicht: (2025)
von: Wang, Xinsheng, et al.
Veröffentlicht: (2025)
U-GIFT: Uncertainty-Guided Firewall for Toxic Speech in Few-Shot Scenario
von: Song, Jiaxin, et al.
Veröffentlicht: (2025)
von: Song, Jiaxin, et al.
Veröffentlicht: (2025)
Zero-Shot Recognition of Dysarthric Speech Using Commercial Automatic Speech Recognition and Multimodal Large Language Models
von: Alsayegh, Ali, et al.
Veröffentlicht: (2025)
von: Alsayegh, Ali, et al.
Veröffentlicht: (2025)
CosyEdit: Unlocking End-to-End Speech Editing Capability from Zero-Shot Text-to-Speech Models
von: Chen, Junyang, et al.
Veröffentlicht: (2026)
von: Chen, Junyang, et al.
Veröffentlicht: (2026)
DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis
von: Tian, Wenjie, et al.
Veröffentlicht: (2025)
von: Tian, Wenjie, et al.
Veröffentlicht: (2025)
HALL-E: Hierarchical Neural Codec Language Model for Minute-Long Zero-Shot Text-to-Speech Synthesis
von: Nishimura, Yuto, et al.
Veröffentlicht: (2024)
von: Nishimura, Yuto, et al.
Veröffentlicht: (2024)
Zero-Shot Text-to-Speech for Vietnamese
von: Vu, Thi, et al.
Veröffentlicht: (2025)
von: Vu, Thi, et al.
Veröffentlicht: (2025)
Parallel Synthesis for Autoregressive Speech Generation
von: Hsu, Po-chun, et al.
Veröffentlicht: (2022)
von: Hsu, Po-chun, et al.
Veröffentlicht: (2022)
VM-UNSSOR: Unsupervised Neural Speech Separation Enhanced by Higher-SNR Virtual Microphone Arrays
von: He, Shulin, et al.
Veröffentlicht: (2025)
von: He, Shulin, et al.
Veröffentlicht: (2025)
Robust Zero-Shot Text-to-Speech Synthesis with Reverse Inference Optimization
von: Hu, Yuchen, et al.
Veröffentlicht: (2024)
von: Hu, Yuchen, et al.
Veröffentlicht: (2024)
Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis
von: Ye, Zhen, et al.
Veröffentlicht: (2025)
von: Ye, Zhen, et al.
Veröffentlicht: (2025)
Llasa+: Free Lunch for Accelerated and Streaming Llama-Based Speech Synthesis
von: Tian, Wenjie, et al.
Veröffentlicht: (2025)
von: Tian, Wenjie, et al.
Veröffentlicht: (2025)
SSR-Speech: Towards Stable, Safe and Robust Zero-shot Text-based Speech Editing and Synthesis
von: Wang, Helin, et al.
Veröffentlicht: (2024)
von: Wang, Helin, et al.
Veröffentlicht: (2024)
SICRN: Advancing Speech Enhancement through State Space Model and Inplace Convolution Techniques
von: Zhao, Changjiang, et al.
Veröffentlicht: (2024)
von: Zhao, Changjiang, et al.
Veröffentlicht: (2024)
VALL-E R: Robust and Efficient Zero-Shot Text-to-Speech Synthesis via Monotonic Alignment
von: Han, Bing, et al.
Veröffentlicht: (2024)
von: Han, Bing, et al.
Veröffentlicht: (2024)
Enabling Beam Search for Language Model-Based Text-to-Speech Synthesis
von: Tu, Zehai, et al.
Veröffentlicht: (2024)
von: Tu, Zehai, et al.
Veröffentlicht: (2024)
ZMM-TTS: Zero-shot Multilingual and Multispeaker Speech Synthesis Conditioned on Self-supervised Discrete Speech Representations
von: Gong, Cheng, et al.
Veröffentlicht: (2023)
von: Gong, Cheng, et al.
Veröffentlicht: (2023)
Inference-time Scaling for Diffusion-based Audio Super-resolution
von: Jin, Yizhu, et al.
Veröffentlicht: (2025)
von: Jin, Yizhu, et al.
Veröffentlicht: (2025)
Multimodal Zero-Shot Framework for Deepfake Hate Speech Detection in Low-Resource Languages
von: Ranjan, Rishabh, et al.
Veröffentlicht: (2025)
von: Ranjan, Rishabh, et al.
Veröffentlicht: (2025)
ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching
von: Zhu, Han, et al.
Veröffentlicht: (2025)
von: Zhu, Han, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models
von: Ju, Zeqian, et al.
Veröffentlicht: (2024) -
Mega-TTS 2: Boosting Prompting Mechanisms for Zero-Shot Speech Synthesis
von: Jiang, Ziyue, et al.
Veröffentlicht: (2023) -
EmoSpeech: A Corpus of Emotionally Rich and Contextually Detailed Speech Annotations
von: Bian, Weizhen, et al.
Veröffentlicht: (2024) -
Towards Expressive Zero-Shot Speech Synthesis with Hierarchical Prosody Modeling
von: Jiang, Yuepeng, et al.
Veröffentlicht: (2024) -
Zero-Shot Audio Captioning Using Soft and Hard Prompts
von: Zhang, Yiming, et al.
Veröffentlicht: (2024)