GenVC: Self-Supervised Zero-Shot Voice Conversion

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Cai, Zexin, Xinyuan, Henry Li, Garg, Ashi, García-Perera, Leibny Paola, Duh, Kevin, Khudanpur, Sanjeev, Wiesner, Matthew, Andrews, Nicholas
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866909745088036864
author Cai, Zexin
Xinyuan, Henry Li
Garg, Ashi
García-Perera, Leibny Paola
Duh, Kevin
Khudanpur, Sanjeev
Wiesner, Matthew
Andrews, Nicholas
author_facet Cai, Zexin
Xinyuan, Henry Li
Garg, Ashi
García-Perera, Leibny Paola
Duh, Kevin
Khudanpur, Sanjeev
Wiesner, Matthew
Andrews, Nicholas
contents Most current zero-shot voice conversion methods rely on externally supervised components, particularly speaker encoders, for training. To explore alternatives that eliminate this dependency, this paper introduces GenVC, a novel framework that disentangles speaker identity and linguistic content from speech signals in a self-supervised manner. GenVC leverages speech tokenizers and an autoregressive, Transformer-based language model as its backbone for speech generation. This design supports large-scale training while enhancing both source speaker privacy protection and target speaker cloning fidelity. Experimental results demonstrate that GenVC achieves notably higher speaker similarity, with naturalness on par with leading zero-shot approaches. Moreover, due to its autoregressive formulation, GenVC introduces flexibility in temporal alignment, reducing the preservation of source prosody and speaker-specific traits, and making it highly effective for voice anonymization.
format Preprint
id arxiv_https___arxiv_org_abs_2502_04519
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle GenVC: Self-Supervised Zero-Shot Voice Conversion
Cai, Zexin
Xinyuan, Henry Li
Garg, Ashi
García-Perera, Leibny Paola
Duh, Kevin
Khudanpur, Sanjeev
Wiesner, Matthew
Andrews, Nicholas
Audio and Speech Processing
Machine Learning
Most current zero-shot voice conversion methods rely on externally supervised components, particularly speaker encoders, for training. To explore alternatives that eliminate this dependency, this paper introduces GenVC, a novel framework that disentangles speaker identity and linguistic content from speech signals in a self-supervised manner. GenVC leverages speech tokenizers and an autoregressive, Transformer-based language model as its backbone for speech generation. This design supports large-scale training while enhancing both source speaker privacy protection and target speaker cloning fidelity. Experimental results demonstrate that GenVC achieves notably higher speaker similarity, with naturalness on par with leading zero-shot approaches. Moreover, due to its autoregressive formulation, GenVC introduces flexibility in temporal alignment, reducing the preservation of source prosody and speaker-specific traits, and making it highly effective for voice anonymization.
title GenVC: Self-Supervised Zero-Shot Voice Conversion
topic Audio and Speech Processing
Machine Learning
url https://arxiv.org/abs/2502.04519