PseudoVC: Improving One-shot Voice Conversion with Pseudo Paired Data

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cao, Songjun, Wu, Qinghua, Chen, Jie, Li, Jin, Ma, Long
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912408817106944
author Cao, Songjun
Wu, Qinghua
Chen, Jie
Li, Jin
Ma, Long
author_facet Cao, Songjun
Wu, Qinghua
Chen, Jie
Li, Jin
Ma, Long
contents As parallel training data is scarce for one-shot voice conversion (VC) tasks, waveform reconstruction is typically performed by various VC systems. A typical one-shot VC system comprises a content encoder and a speaker encoder. However, two types of mismatches arise: one for the inputs to the content encoder during training and inference, and another for the inputs to the speaker encoder. To address these mismatches, we propose a novel VC training method called \textit{PseudoVC} in this paper. First, we introduce an innovative information perturbation approach named \textit{Pseudo Conversion} to tackle the first mismatch problem. This approach leverages pretrained VC models to convert the source utterance into a perturbed utterance, which is fed into the content encoder during training. Second, we propose an approach termed \textit{Speaker Sampling} to resolve the second mismatch problem, which will substitute the input to the speaker encoder by another utterance from the same speaker during training. Experimental results demonstrate that our proposed \textit{Pseudo Conversion} outperforms previous information perturbation methods, and the overall \textit{PseudoVC} method surpasses publicly available VC models. Audio examples are available.
format Preprint
id arxiv_https___arxiv_org_abs_2506_01039
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle PseudoVC: Improving One-shot Voice Conversion with Pseudo Paired Data
Cao, Songjun
Wu, Qinghua
Chen, Jie
Li, Jin
Ma, Long
Audio and Speech Processing
Sound
As parallel training data is scarce for one-shot voice conversion (VC) tasks, waveform reconstruction is typically performed by various VC systems. A typical one-shot VC system comprises a content encoder and a speaker encoder. However, two types of mismatches arise: one for the inputs to the content encoder during training and inference, and another for the inputs to the speaker encoder. To address these mismatches, we propose a novel VC training method called \textit{PseudoVC} in this paper. First, we introduce an innovative information perturbation approach named \textit{Pseudo Conversion} to tackle the first mismatch problem. This approach leverages pretrained VC models to convert the source utterance into a perturbed utterance, which is fed into the content encoder during training. Second, we propose an approach termed \textit{Speaker Sampling} to resolve the second mismatch problem, which will substitute the input to the speaker encoder by another utterance from the same speaker during training. Experimental results demonstrate that our proposed \textit{Pseudo Conversion} outperforms previous information perturbation methods, and the overall \textit{PseudoVC} method surpasses publicly available VC models. Audio examples are available.
title PseudoVC: Improving One-shot Voice Conversion with Pseudo Paired Data
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2506.01039