X-VC: Zero-shot Streaming Voice Conversion in Codec Space

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zheng, Qixi, Zhao, Yuxiang, Wang, Tianrui, Chen, Wenxi, Xu, Kele, Li, Yikang, Chen, Qinyuan, Qiu, Xipeng, Yu, Kai, Chen, Xie
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917427055427584
author Zheng, Qixi
Zhao, Yuxiang
Wang, Tianrui
Chen, Wenxi
Xu, Kele
Li, Yikang
Chen, Qinyuan
Qiu, Xipeng
Yu, Kai
Chen, Xie
author_facet Zheng, Qixi
Zhao, Yuxiang
Wang, Tianrui
Chen, Wenxi
Xu, Kele
Li, Yikang
Chen, Qinyuan
Qiu, Xipeng
Yu, Kai
Chen, Xie
contents Zero-shot voice conversion (VC) aims to convert a source utterance into the voice of an unseen target speaker while preserving its linguistic content. Although recent systems have improved conversion quality, building zero-shot VC systems for interactive scenarios remains challenging because high-fidelity speaker transfer and low-latency streaming inference are difficult to achieve simultaneously. In this work, we present X-VC, a zero-shot streaming VC system that performs one-step conversion in the latent space of a pretrained neural codec. X-VC uses a dual-conditioning acoustic converter that jointly models source codec latents and frame-level acoustic conditions derived from target reference speech, while injecting utterance-level target speaker information through adaptive normalization. To reduce the mismatch between training and inference, we train the model with generated paired data and a role-assignment strategy that combines standard, reconstruction, and reversed modes. For streaming inference, we further adopt a chunkwise inference scheme with overlap smoothing that is aligned with the segment-based training paradigm of the codec. Experiments on Seed-TTS-Eval show that X-VC achieves the best streaming WER in both English and Chinese, strong speaker similarity in same-language and cross-lingual settings, and substantially lower offline real-time factor than the compared baselines. These results suggest that codec-space one-step conversion is a practical approach for building high-quality low-latency zero-shot VC systems. Our audio samples, code and checkpoints are released at https://github.com/Jerrister/X-VC.
format Preprint
id arxiv_https___arxiv_org_abs_2604_12456
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle X-VC: Zero-shot Streaming Voice Conversion in Codec Space
Zheng, Qixi
Zhao, Yuxiang
Wang, Tianrui
Chen, Wenxi
Xu, Kele
Li, Yikang
Chen, Qinyuan
Qiu, Xipeng
Yu, Kai
Chen, Xie
Audio and Speech Processing
Artificial Intelligence
Zero-shot voice conversion (VC) aims to convert a source utterance into the voice of an unseen target speaker while preserving its linguistic content. Although recent systems have improved conversion quality, building zero-shot VC systems for interactive scenarios remains challenging because high-fidelity speaker transfer and low-latency streaming inference are difficult to achieve simultaneously. In this work, we present X-VC, a zero-shot streaming VC system that performs one-step conversion in the latent space of a pretrained neural codec. X-VC uses a dual-conditioning acoustic converter that jointly models source codec latents and frame-level acoustic conditions derived from target reference speech, while injecting utterance-level target speaker information through adaptive normalization. To reduce the mismatch between training and inference, we train the model with generated paired data and a role-assignment strategy that combines standard, reconstruction, and reversed modes. For streaming inference, we further adopt a chunkwise inference scheme with overlap smoothing that is aligned with the segment-based training paradigm of the codec. Experiments on Seed-TTS-Eval show that X-VC achieves the best streaming WER in both English and Chinese, strong speaker similarity in same-language and cross-lingual settings, and substantially lower offline real-time factor than the compared baselines. These results suggest that codec-space one-step conversion is a practical approach for building high-quality low-latency zero-shot VC systems. Our audio samples, code and checkpoints are released at https://github.com/Jerrister/X-VC.
title X-VC: Zero-shot Streaming Voice Conversion in Codec Space
topic Audio and Speech Processing
Artificial Intelligence
url https://arxiv.org/abs/2604.12456