Generating Novel and Realistic Speakers for Voice Conversion

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Meiying Melissa, Wang, Zhenyu, Duan, Zhiyao
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909896228732928
author Chen, Meiying Melissa
Wang, Zhenyu
Duan, Zhiyao
author_facet Chen, Meiying Melissa
Wang, Zhenyu
Duan, Zhiyao
contents Voice conversion models modify timbre while preserving paralinguistic features, enabling applications like dubbing and identity protection. However, most VC systems require access to target utterances, limiting their use when target data is unavailable or when users desire conversion to entirely novel, unseen voices. To address this, we introduce a lightweight method SpeakerVAE to generate novel speakers for VC. Our approach uses a deep hierarchical variational autoencoder to model the speaker timbre space. By sampling from the trained model, we generate novel speaker representations for voice synthesis in a VC pipeline. The proposed method is a flexible plug-in module compatible with various VC models, without co-training or fine-tuning of the base VC system. We evaluated our approach with state-of-the-art VC models: FACodec and CosyVoice2. The results demonstrate that our method successfully generates novel, unseen speakers with quality comparable to that of the training speakers.
format Preprint
id arxiv_https___arxiv_org_abs_2511_07135
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Generating Novel and Realistic Speakers for Voice Conversion
Chen, Meiying Melissa
Wang, Zhenyu
Duan, Zhiyao
Sound
Audio and Speech Processing
Voice conversion models modify timbre while preserving paralinguistic features, enabling applications like dubbing and identity protection. However, most VC systems require access to target utterances, limiting their use when target data is unavailable or when users desire conversion to entirely novel, unseen voices. To address this, we introduce a lightweight method SpeakerVAE to generate novel speakers for VC. Our approach uses a deep hierarchical variational autoencoder to model the speaker timbre space. By sampling from the trained model, we generate novel speaker representations for voice synthesis in a VC pipeline. The proposed method is a flexible plug-in module compatible with various VC models, without co-training or fine-tuning of the base VC system. We evaluated our approach with state-of-the-art VC models: FACodec and CosyVoice2. The results demonstrate that our method successfully generates novel, unseen speakers with quality comparable to that of the training speakers.
title Generating Novel and Realistic Speakers for Voice Conversion
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2511.07135