MaskVCT: Masked Voice Codec Transformer for Zero-Shot Voice Conversion With Increased Controllability via Multiple Guidances

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Lee, Junhyeok, Wang, Helin, Guan, Yaohan, Thebaud, Thomas, Moro-Velazquez, Laureano, Villalba, Jesús, Dehak, Najim
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914322659147776
author Lee, Junhyeok
Wang, Helin
Guan, Yaohan
Thebaud, Thomas
Moro-Velazquez, Laureano
Villalba, Jesús
Dehak, Najim
author_facet Lee, Junhyeok
Wang, Helin
Guan, Yaohan
Thebaud, Thomas
Moro-Velazquez, Laureano
Villalba, Jesús
Dehak, Najim
contents We introduce MaskVCT, a zero-shot voice conversion (VC) model that offers multi-factor controllability through multiple classifier-free guidances (CFGs). While previous VC models rely on a fixed conditioning scheme, MaskVCT integrates diverse conditions in a single model. To further enhance robustness and control, the model can leverage continuous or quantized linguistic features to enhance intelligibility and speaker similarity, and can use or omit pitch contour to control prosody. These choices allow users to seamlessly balance speaker identity, linguistic content, and prosodic factors in a zero-shot VC setting. Extensive experiments demonstrate that MaskVCT achieves the best target speaker and accent similarities while obtaining competitive word and character error rates compared to existing baselines. Audio samples are available at https://maskvct.github.io/.
format Preprint
id arxiv_https___arxiv_org_abs_2509_17143
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MaskVCT: Masked Voice Codec Transformer for Zero-Shot Voice Conversion With Increased Controllability via Multiple Guidances
Lee, Junhyeok
Wang, Helin
Guan, Yaohan
Thebaud, Thomas
Moro-Velazquez, Laureano
Villalba, Jesús
Dehak, Najim
Audio and Speech Processing
Artificial Intelligence
We introduce MaskVCT, a zero-shot voice conversion (VC) model that offers multi-factor controllability through multiple classifier-free guidances (CFGs). While previous VC models rely on a fixed conditioning scheme, MaskVCT integrates diverse conditions in a single model. To further enhance robustness and control, the model can leverage continuous or quantized linguistic features to enhance intelligibility and speaker similarity, and can use or omit pitch contour to control prosody. These choices allow users to seamlessly balance speaker identity, linguistic content, and prosodic factors in a zero-shot VC setting. Extensive experiments demonstrate that MaskVCT achieves the best target speaker and accent similarities while obtaining competitive word and character error rates compared to existing baselines. Audio samples are available at https://maskvct.github.io/.
title MaskVCT: Masked Voice Codec Transformer for Zero-Shot Voice Conversion With Increased Controllability via Multiple Guidances
topic Audio and Speech Processing
Artificial Intelligence
url https://arxiv.org/abs/2509.17143