X-Voice: Enabling Everyone to Speak 30 Languages via Zero-Shot Cross-Lingual Voice Cloning

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Xu, Rixi, Liu, Qingyu, Li, Haitao, Chen, Yushen, Niu, Zhikang, Yang, Yunting, Zhao, Jian, Li, Ke, Sisman, Berrak, Cheng, Qinyuan, Qiu, Xipeng, Yu, Kai, Chen, Xie
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915997266477056
author Xu, Rixi
Liu, Qingyu
Li, Haitao
Chen, Yushen
Niu, Zhikang
Yang, Yunting
Zhao, Jian
Li, Ke
Sisman, Berrak
Cheng, Qinyuan
Qiu, Xipeng
Yu, Kai
Chen, Xie
author_facet Xu, Rixi
Liu, Qingyu
Li, Haitao
Chen, Yushen
Niu, Zhikang
Yang, Yunting
Zhao, Jian
Li, Ke
Sisman, Berrak
Cheng, Qinyuan
Qiu, Xipeng
Yu, Kai
Chen, Xie
contents In this paper, we present X-Voice, a 0.4B multilingual zero-shot voice cloning model that clones arbitrary voices and enables everyone to speak 30 languages. X-Voice is trained on a 420K-hour multilingual corpus using the International Phonetic Alphabet (IPA) as a unified representation. To eliminate the reliance on prompt text without complex preprocessing like forced alignment, we design a two-stage training paradigm. In Stage 1, we establish X-Voice$_{\text{s1}}$ through standard conditional flow-matching training and use it to synthesize 10K hours of speaker-consistent segments as audio prompts. In Stage 2, we fine-tune on these audio pairs with prompt text masked to derive X-Voice$_{\text{s2}}$, which enables zero-shot voice cloning without requiring transcripts of audio prompts. Architecturally, we extend F5-TTS by implementing a dual-level injection of language identifiers and decoupling and scheduling of Classifier-Free Guidance to facilitate multilingual speech synthesis. Subjective and objective evaluation results demonstrate that X-Voice outperforms existing flow-matching based multilingual systems like LEMAS-TTS and achieves zero-shot cross-lingual cloning capabilities comparable to billion-scale models such as Qwen3-TTS. To facilitate research transparency and community advancement, we open-source all related resources.
format Preprint
id arxiv_https___arxiv_org_abs_2605_05611
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle X-Voice: Enabling Everyone to Speak 30 Languages via Zero-Shot Cross-Lingual Voice Cloning
Xu, Rixi
Liu, Qingyu
Li, Haitao
Chen, Yushen
Niu, Zhikang
Yang, Yunting
Zhao, Jian
Li, Ke
Sisman, Berrak
Cheng, Qinyuan
Qiu, Xipeng
Yu, Kai
Chen, Xie
Sound
Artificial Intelligence
Audio and Speech Processing
In this paper, we present X-Voice, a 0.4B multilingual zero-shot voice cloning model that clones arbitrary voices and enables everyone to speak 30 languages. X-Voice is trained on a 420K-hour multilingual corpus using the International Phonetic Alphabet (IPA) as a unified representation. To eliminate the reliance on prompt text without complex preprocessing like forced alignment, we design a two-stage training paradigm. In Stage 1, we establish X-Voice$_{\text{s1}}$ through standard conditional flow-matching training and use it to synthesize 10K hours of speaker-consistent segments as audio prompts. In Stage 2, we fine-tune on these audio pairs with prompt text masked to derive X-Voice$_{\text{s2}}$, which enables zero-shot voice cloning without requiring transcripts of audio prompts. Architecturally, we extend F5-TTS by implementing a dual-level injection of language identifiers and decoupling and scheduling of Classifier-Free Guidance to facilitate multilingual speech synthesis. Subjective and objective evaluation results demonstrate that X-Voice outperforms existing flow-matching based multilingual systems like LEMAS-TTS and achieves zero-shot cross-lingual cloning capabilities comparable to billion-scale models such as Qwen3-TTS. To facilitate research transparency and community advancement, we open-source all related resources.
title X-Voice: Enabling Everyone to Speak 30 Languages via Zero-Shot Cross-Lingual Voice Cloning
topic Sound
Artificial Intelligence
Audio and Speech Processing
url https://arxiv.org/abs/2605.05611