Discl-VC: Disentangled Discrete Tokens and In-Context Learning for Controllable Zero-Shot Voice Conversion

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Kaidi, Guan, Wenhao, Jiang, Ziyue, Huang, Hukai, Chen, Peijie, Wu, Weijie, Hong, Qingyang, Li, Lin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918039369285632
author Wang, Kaidi
Guan, Wenhao
Jiang, Ziyue
Huang, Hukai
Chen, Peijie
Wu, Weijie
Hong, Qingyang
Li, Lin
author_facet Wang, Kaidi
Guan, Wenhao
Jiang, Ziyue
Huang, Hukai
Chen, Peijie
Wu, Weijie
Hong, Qingyang
Li, Lin
contents Currently, zero-shot voice conversion systems are capable of synthesizing the voice of unseen speakers. However, most existing approaches struggle to accurately replicate the speaking style of the source speaker or mimic the distinctive speaking style of the target speaker, thereby limiting the controllability of voice conversion. In this work, we propose Discl-VC, a novel voice conversion framework that disentangles content and prosody information from self-supervised speech representations and synthesizes the target speaker's voice through in-context learning with a flow matching transformer. To enable precise control over the prosody of generated speech, we introduce a mask generative transformer that predicts discrete prosody tokens in a non-autoregressive manner based on prompts. Experimental results demonstrate the superior performance of Discl-VC in zero-shot voice conversion and its remarkable accuracy in prosody control for synthesized speech.
format Preprint
id arxiv_https___arxiv_org_abs_2505_24291
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Discl-VC: Disentangled Discrete Tokens and In-Context Learning for Controllable Zero-Shot Voice Conversion
Wang, Kaidi
Guan, Wenhao
Jiang, Ziyue
Huang, Hukai
Chen, Peijie
Wu, Weijie
Hong, Qingyang
Li, Lin
Sound
Artificial Intelligence
Audio and Speech Processing
Currently, zero-shot voice conversion systems are capable of synthesizing the voice of unseen speakers. However, most existing approaches struggle to accurately replicate the speaking style of the source speaker or mimic the distinctive speaking style of the target speaker, thereby limiting the controllability of voice conversion. In this work, we propose Discl-VC, a novel voice conversion framework that disentangles content and prosody information from self-supervised speech representations and synthesizes the target speaker's voice through in-context learning with a flow matching transformer. To enable precise control over the prosody of generated speech, we introduce a mask generative transformer that predicts discrete prosody tokens in a non-autoregressive manner based on prompts. Experimental results demonstrate the superior performance of Discl-VC in zero-shot voice conversion and its remarkable accuracy in prosody control for synthesized speech.
title Discl-VC: Disentangled Discrete Tokens and In-Context Learning for Controllable Zero-Shot Voice Conversion
topic Sound
Artificial Intelligence
Audio and Speech Processing
url https://arxiv.org/abs/2505.24291