VoiceCraft-X: Unifying Multilingual, Voice-Cloning Speech Synthesis and Speech Editing
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912712171192320 |
|---|---|
| author | Zheng, Zhisheng Peng, Puyuan Diwan, Anuj Huynh, Cong Phuoc Sun, Xiaohang Liu, Zhu Bhat, Vimal Harwath, David |
| author_facet | Zheng, Zhisheng Peng, Puyuan Diwan, Anuj Huynh, Cong Phuoc Sun, Xiaohang Liu, Zhu Bhat, Vimal Harwath, David |
| contents | We introduce VoiceCraft-X, an autoregressive neural codec language model which unifies multilingual speech editing and zero-shot Text-to-Speech (TTS) synthesis across 11 languages: English, Mandarin, Korean, Japanese, Spanish, French, German, Dutch, Italian, Portuguese, and Polish. VoiceCraft-X utilizes the Qwen3 large language model for phoneme-free cross-lingual text processing and a novel token reordering mechanism with time-aligned text and speech tokens to handle both tasks as a single sequence generation problem. The model generates high-quality, natural-sounding speech, seamlessly creating new audio or editing existing recordings within one framework. VoiceCraft-X shows robust performance in diverse linguistic settings, even with limited per-language data, underscoring the power of unified autoregressive approaches for advancing complex, real-world multilingual speech applications. Audio samples are available at https://zhishengzheng.com/voicecraft-x/. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2511_12347 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | VoiceCraft-X: Unifying Multilingual, Voice-Cloning Speech Synthesis and Speech Editing Zheng, Zhisheng Peng, Puyuan Diwan, Anuj Huynh, Cong Phuoc Sun, Xiaohang Liu, Zhu Bhat, Vimal Harwath, David Audio and Speech Processing Computation and Language Sound We introduce VoiceCraft-X, an autoregressive neural codec language model which unifies multilingual speech editing and zero-shot Text-to-Speech (TTS) synthesis across 11 languages: English, Mandarin, Korean, Japanese, Spanish, French, German, Dutch, Italian, Portuguese, and Polish. VoiceCraft-X utilizes the Qwen3 large language model for phoneme-free cross-lingual text processing and a novel token reordering mechanism with time-aligned text and speech tokens to handle both tasks as a single sequence generation problem. The model generates high-quality, natural-sounding speech, seamlessly creating new audio or editing existing recordings within one framework. VoiceCraft-X shows robust performance in diverse linguistic settings, even with limited per-language data, underscoring the power of unified autoregressive approaches for advancing complex, real-world multilingual speech applications. Audio samples are available at https://zhishengzheng.com/voicecraft-x/. |
| title | VoiceCraft-X: Unifying Multilingual, Voice-Cloning Speech Synthesis and Speech Editing |
| topic | Audio and Speech Processing Computation and Language Sound |
| url | https://arxiv.org/abs/2511.12347 |