GenVC: Self-Supervised Zero-Shot Voice Conversion
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866909745088036864 |
|---|---|
| author | Cai, Zexin Xinyuan, Henry Li Garg, Ashi García-Perera, Leibny Paola Duh, Kevin Khudanpur, Sanjeev Wiesner, Matthew Andrews, Nicholas |
| author_facet | Cai, Zexin Xinyuan, Henry Li Garg, Ashi García-Perera, Leibny Paola Duh, Kevin Khudanpur, Sanjeev Wiesner, Matthew Andrews, Nicholas |
| contents | Most current zero-shot voice conversion methods rely on externally supervised components, particularly speaker encoders, for training. To explore alternatives that eliminate this dependency, this paper introduces GenVC, a novel framework that disentangles speaker identity and linguistic content from speech signals in a self-supervised manner. GenVC leverages speech tokenizers and an autoregressive, Transformer-based language model as its backbone for speech generation. This design supports large-scale training while enhancing both source speaker privacy protection and target speaker cloning fidelity. Experimental results demonstrate that GenVC achieves notably higher speaker similarity, with naturalness on par with leading zero-shot approaches. Moreover, due to its autoregressive formulation, GenVC introduces flexibility in temporal alignment, reducing the preservation of source prosody and speaker-specific traits, and making it highly effective for voice anonymization. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2502_04519 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | GenVC: Self-Supervised Zero-Shot Voice Conversion Cai, Zexin Xinyuan, Henry Li Garg, Ashi García-Perera, Leibny Paola Duh, Kevin Khudanpur, Sanjeev Wiesner, Matthew Andrews, Nicholas Audio and Speech Processing Machine Learning Most current zero-shot voice conversion methods rely on externally supervised components, particularly speaker encoders, for training. To explore alternatives that eliminate this dependency, this paper introduces GenVC, a novel framework that disentangles speaker identity and linguistic content from speech signals in a self-supervised manner. GenVC leverages speech tokenizers and an autoregressive, Transformer-based language model as its backbone for speech generation. This design supports large-scale training while enhancing both source speaker privacy protection and target speaker cloning fidelity. Experimental results demonstrate that GenVC achieves notably higher speaker similarity, with naturalness on par with leading zero-shot approaches. Moreover, due to its autoregressive formulation, GenVC introduces flexibility in temporal alignment, reducing the preservation of source prosody and speaker-specific traits, and making it highly effective for voice anonymization. |
| title | GenVC: Self-Supervised Zero-Shot Voice Conversion |
| topic | Audio and Speech Processing Machine Learning |
| url | https://arxiv.org/abs/2502.04519 |