VoiceCraft-X: Unifying Multilingual, Voice-Cloning Speech Synthesis and Speech Editing

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zheng, Zhisheng, Peng, Puyuan, Diwan, Anuj, Huynh, Cong Phuoc, Sun, Xiaohang, Liu, Zhu, Bhat, Vimal, Harwath, David
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912712171192320
author Zheng, Zhisheng
Peng, Puyuan
Diwan, Anuj
Huynh, Cong Phuoc
Sun, Xiaohang
Liu, Zhu
Bhat, Vimal
Harwath, David
author_facet Zheng, Zhisheng
Peng, Puyuan
Diwan, Anuj
Huynh, Cong Phuoc
Sun, Xiaohang
Liu, Zhu
Bhat, Vimal
Harwath, David
contents We introduce VoiceCraft-X, an autoregressive neural codec language model which unifies multilingual speech editing and zero-shot Text-to-Speech (TTS) synthesis across 11 languages: English, Mandarin, Korean, Japanese, Spanish, French, German, Dutch, Italian, Portuguese, and Polish. VoiceCraft-X utilizes the Qwen3 large language model for phoneme-free cross-lingual text processing and a novel token reordering mechanism with time-aligned text and speech tokens to handle both tasks as a single sequence generation problem. The model generates high-quality, natural-sounding speech, seamlessly creating new audio or editing existing recordings within one framework. VoiceCraft-X shows robust performance in diverse linguistic settings, even with limited per-language data, underscoring the power of unified autoregressive approaches for advancing complex, real-world multilingual speech applications. Audio samples are available at https://zhishengzheng.com/voicecraft-x/.
format Preprint
id arxiv_https___arxiv_org_abs_2511_12347
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle VoiceCraft-X: Unifying Multilingual, Voice-Cloning Speech Synthesis and Speech Editing
Zheng, Zhisheng
Peng, Puyuan
Diwan, Anuj
Huynh, Cong Phuoc
Sun, Xiaohang
Liu, Zhu
Bhat, Vimal
Harwath, David
Audio and Speech Processing
Computation and Language
Sound
We introduce VoiceCraft-X, an autoregressive neural codec language model which unifies multilingual speech editing and zero-shot Text-to-Speech (TTS) synthesis across 11 languages: English, Mandarin, Korean, Japanese, Spanish, French, German, Dutch, Italian, Portuguese, and Polish. VoiceCraft-X utilizes the Qwen3 large language model for phoneme-free cross-lingual text processing and a novel token reordering mechanism with time-aligned text and speech tokens to handle both tasks as a single sequence generation problem. The model generates high-quality, natural-sounding speech, seamlessly creating new audio or editing existing recordings within one framework. VoiceCraft-X shows robust performance in diverse linguistic settings, even with limited per-language data, underscoring the power of unified autoregressive approaches for advancing complex, real-world multilingual speech applications. Audio samples are available at https://zhishengzheng.com/voicecraft-x/.
title VoiceCraft-X: Unifying Multilingual, Voice-Cloning Speech Synthesis and Speech Editing
topic Audio and Speech Processing
Computation and Language
Sound
url https://arxiv.org/abs/2511.12347