VoiceCraft: Zero-Shot Speech Editing and Text-to-Speech in the Wild
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Peng, Puyuan, Huang, Po-Yao, Li, Shang-Wen, Mohamed, Abdelrahman, Harwath, David |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
VoiceCraft-X: Unifying Multilingual, Voice-Cloning Speech Synthesis and Speech Editing
von: Zheng, Zhisheng, et al.
Veröffentlicht: (2025)
von: Zheng, Zhisheng, et al.
Veröffentlicht: (2025)
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation
von: Peng, Puyuan, et al.
Veröffentlicht: (2025)
von: Peng, Puyuan, et al.
Veröffentlicht: (2025)
Zero-Shot Text-to-Speech for Vietnamese
von: Vu, Thi, et al.
Veröffentlicht: (2025)
von: Vu, Thi, et al.
Veröffentlicht: (2025)
Interface Design for Self-Supervised Speech Models
von: Shih, Yi-Jen, et al.
Veröffentlicht: (2024)
von: Shih, Yi-Jen, et al.
Veröffentlicht: (2024)
SF-Speech: Straightened Flow for Zero-Shot Voice Clone
von: Li, Xuyuan, et al.
Veröffentlicht: (2024)
von: Li, Xuyuan, et al.
Veröffentlicht: (2024)
ParaSpeechCLAP: A Dual-Encoder Speech-Text Model for Rich Stylistic Language-Audio Pretraining
von: Diwan, Anuj, et al.
Veröffentlicht: (2026)
von: Diwan, Anuj, et al.
Veröffentlicht: (2026)
Towards Zero-Shot Text-To-Speech for Arabic Dialects
von: Doan, Khai Duy, et al.
Veröffentlicht: (2024)
von: Doan, Khai Duy, et al.
Veröffentlicht: (2024)
VoiceShop: A Unified Speech-to-Speech Framework for Identity-Preserving Zero-Shot Voice Editing
von: Anastassiou, Philip, et al.
Veröffentlicht: (2024)
von: Anastassiou, Philip, et al.
Veröffentlicht: (2024)
ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching
von: Zhu, Han, et al.
Veröffentlicht: (2025)
von: Zhu, Han, et al.
Veröffentlicht: (2025)
SpeechCLIP+: Self-supervised multi-task representation learning for speech via CLIP and speech-image data
von: Wang, Hsuan-Fu, et al.
Veröffentlicht: (2024)
von: Wang, Hsuan-Fu, et al.
Veröffentlicht: (2024)
VoiceCraft-Dub: Automated Video Dubbing with Neural Codec Language Models
von: Sung-Bin, Kim, et al.
Veröffentlicht: (2025)
von: Sung-Bin, Kim, et al.
Veröffentlicht: (2025)
CosyEdit: Unlocking End-to-End Speech Editing Capability from Zero-Shot Text-to-Speech Models
von: Chen, Junyang, et al.
Veröffentlicht: (2026)
von: Chen, Junyang, et al.
Veröffentlicht: (2026)
XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model
von: Casanova, Edresson, et al.
Veröffentlicht: (2024)
von: Casanova, Edresson, et al.
Veröffentlicht: (2024)
Robust Zero-Shot Text-to-Speech Synthesis with Reverse Inference Optimization
von: Hu, Yuchen, et al.
Veröffentlicht: (2024)
von: Hu, Yuchen, et al.
Veröffentlicht: (2024)
Probing the Robustness Properties of Neural Speech Codecs
von: Tseng, Wei-Cheng, et al.
Veröffentlicht: (2025)
von: Tseng, Wei-Cheng, et al.
Veröffentlicht: (2025)
Zero-Shot Text-to-Speech from Continuous Text Streams
von: Dang, Trung, et al.
Veröffentlicht: (2024)
von: Dang, Trung, et al.
Veröffentlicht: (2024)
Zero Resource Code-switched Speech Benchmark Using Speech Utterance Pairs For Multiple Spoken Languages
von: Huang, Kuan-Po, et al.
Veröffentlicht: (2023)
von: Huang, Kuan-Po, et al.
Veröffentlicht: (2023)
SyllableLM: Learning Coarse Semantic Units for Speech Language Models
von: Baade, Alan, et al.
Veröffentlicht: (2024)
von: Baade, Alan, et al.
Veröffentlicht: (2024)
VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers
von: Chen, Sanyuan, et al.
Veröffentlicht: (2024)
von: Chen, Sanyuan, et al.
Veröffentlicht: (2024)
Enhancing Zero-shot Text-to-Speech Synthesis with Human Feedback
von: Chen, Chen, et al.
Veröffentlicht: (2024)
von: Chen, Chen, et al.
Veröffentlicht: (2024)
Zero Shot Text to Speech Augmentation for Automatic Speech Recognition on Low-Resource Accented Speech Corpora
von: Nespoli, Francesco, et al.
Veröffentlicht: (2024)
von: Nespoli, Francesco, et al.
Veröffentlicht: (2024)
MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder
von: Zhang, Bowen, et al.
Veröffentlicht: (2025)
von: Zhang, Bowen, et al.
Veröffentlicht: (2025)
Self-supervised Speech Models for Word-Level Stuttered Speech Detection
von: Shih, Yi-Jen, et al.
Veröffentlicht: (2024)
von: Shih, Yi-Jen, et al.
Veröffentlicht: (2024)
ML-SUPERB: Multilingual Speech Universal PERformance Benchmark
von: Shi, Jiatong, et al.
Veröffentlicht: (2023)
von: Shi, Jiatong, et al.
Veröffentlicht: (2023)
VALL-E R: Robust and Efficient Zero-Shot Text-to-Speech Synthesis via Monotonic Alignment
von: Han, Bing, et al.
Veröffentlicht: (2024)
von: Han, Bing, et al.
Veröffentlicht: (2024)
Transliterated Zero-Shot Domain Adaptation for Automatic Speech Recognition
von: Zhu, Han, et al.
Veröffentlicht: (2024)
von: Zhu, Han, et al.
Veröffentlicht: (2024)
Advanced Zero-Shot Text-to-Speech for Background Removal and Preservation with Controllable Masked Speech Prediction
von: Zhang, Leying, et al.
Veröffentlicht: (2025)
von: Zhang, Leying, et al.
Veröffentlicht: (2025)
MobileSpeech: A Fast and High-Fidelity Framework for Mobile Zero-Shot Text-to-Speech
von: Ji, Shengpeng, et al.
Veröffentlicht: (2024)
von: Ji, Shengpeng, et al.
Veröffentlicht: (2024)
SSR-Speech: Towards Stable, Safe and Robust Zero-shot Text-based Speech Editing and Synthesis
von: Wang, Helin, et al.
Veröffentlicht: (2024)
von: Wang, Helin, et al.
Veröffentlicht: (2024)
Everyone-Can-Sing: Zero-Shot Singing Voice Synthesis and Conversion with Speech Reference
von: Dai, Shuqi, et al.
Veröffentlicht: (2025)
von: Dai, Shuqi, et al.
Veröffentlicht: (2025)
Scaling Rich Style-Prompted Text-to-Speech Datasets
von: Diwan, Anuj, et al.
Veröffentlicht: (2025)
von: Diwan, Anuj, et al.
Veröffentlicht: (2025)
Word-Level Emotional Expression Control in Zero-Shot Text-to-Speech Synthesis
von: Wang, Tianrui, et al.
Veröffentlicht: (2025)
von: Wang, Tianrui, et al.
Veröffentlicht: (2025)
BAT: Learning to Reason about Spatial Sounds with Large Language Models
von: Zheng, Zhisheng, et al.
Veröffentlicht: (2024)
von: Zheng, Zhisheng, et al.
Veröffentlicht: (2024)
Speech Editing -- a Summary
von: Kässmann, Tobias, et al.
Veröffentlicht: (2024)
von: Kässmann, Tobias, et al.
Veröffentlicht: (2024)
Continuous Speech Tokenizer in Text To Speech
von: Li, Yixing, et al.
Veröffentlicht: (2024)
von: Li, Yixing, et al.
Veröffentlicht: (2024)
Improving Audio Codec-based Zero-Shot Text-to-Speech Synthesis with Multi-Modal Context and Large Language Model
von: Xue, Jinlong, et al.
Veröffentlicht: (2024)
von: Xue, Jinlong, et al.
Veröffentlicht: (2024)
Improvement Speaker Similarity for Zero-Shot Any-to-Any Voice Conversion of Whispered and Regular Speech
von: Avdeeva, Anastasia, et al.
Veröffentlicht: (2024)
von: Avdeeva, Anastasia, et al.
Veröffentlicht: (2024)
Instance-Specific Test-Time Training for Speech Editing in the Wild
von: Kim, Taewoo, et al.
Veröffentlicht: (2025)
von: Kim, Taewoo, et al.
Veröffentlicht: (2025)
Do Not Mimic My Voice: Speaker Identity Unlearning for Zero-Shot Text-to-Speech
von: Kim, Taesoo, et al.
Veröffentlicht: (2025)
von: Kim, Taesoo, et al.
Veröffentlicht: (2025)
Generalized Multilingual Text-to-Speech Generation with Language-Aware Style Adaptation
von: Lou, Haowei, et al.
Veröffentlicht: (2025)
von: Lou, Haowei, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
VoiceCraft-X: Unifying Multilingual, Voice-Cloning Speech Synthesis and Speech Editing
von: Zheng, Zhisheng, et al.
Veröffentlicht: (2025) -
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation
von: Peng, Puyuan, et al.
Veröffentlicht: (2025) -
Zero-Shot Text-to-Speech for Vietnamese
von: Vu, Thi, et al.
Veröffentlicht: (2025) -
Interface Design for Self-Supervised Speech Models
von: Shih, Yi-Jen, et al.
Veröffentlicht: (2024) -
SF-Speech: Straightened Flow for Zero-Shot Voice Clone
von: Li, Xuyuan, et al.
Veröffentlicht: (2024)