Saved in:
Bibliographic Details
Main Authors: Liang, Ziqi, Zhang, Xulong, Liu, Chang, Qu, Xiaoyang, Zhao, Weifeng, Wang, Jianzong
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2501.01861
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909447719223296
author Liang, Ziqi
Zhang, Xulong
Liu, Chang
Qu, Xiaoyang
Zhao, Weifeng
Wang, Jianzong
author_facet Liang, Ziqi
Zhang, Xulong
Liu, Chang
Qu, Xiaoyang
Zhao, Weifeng
Wang, Jianzong
contents Voice Conversion (VC) aims to convert the style of a source speaker, such as timbre and pitch, to the style of any target speaker while preserving the linguistic content. However, the ground truth of the converted speech does not exist in a non-parallel VC scenario, which induces the train-inference mismatch problem. Moreover, existing methods still have an inaccurate pitch and low speaker adaptation quality, there is a significant disparity in pitch between the source and target speaker style domains. As a result, the models tend to generate speech with hoarseness, posing challenges in achieving high-quality voice conversion. In this study, we propose CycleFlow, a novel VC approach that leverages cycle consistency in conditional flow matching (CFM) for speaker timbre adaptation training on non-parallel data. Furthermore, we design a Dual-CFM based on VoiceCFM and PitchCFM to generate speech and improve speaker pitch adaptation quality. Experiments show that our method can significantly improve speaker similarity, generating natural and higher-quality speech.
format Preprint
id arxiv_https___arxiv_org_abs_2501_01861
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CycleFlow: Leveraging Cycle Consistency in Flow Matching for Speaker Style Adaptation
Liang, Ziqi
Zhang, Xulong
Liu, Chang
Qu, Xiaoyang
Zhao, Weifeng
Wang, Jianzong
Sound
Audio and Speech Processing
Voice Conversion (VC) aims to convert the style of a source speaker, such as timbre and pitch, to the style of any target speaker while preserving the linguistic content. However, the ground truth of the converted speech does not exist in a non-parallel VC scenario, which induces the train-inference mismatch problem. Moreover, existing methods still have an inaccurate pitch and low speaker adaptation quality, there is a significant disparity in pitch between the source and target speaker style domains. As a result, the models tend to generate speech with hoarseness, posing challenges in achieving high-quality voice conversion. In this study, we propose CycleFlow, a novel VC approach that leverages cycle consistency in conditional flow matching (CFM) for speaker timbre adaptation training on non-parallel data. Furthermore, we design a Dual-CFM based on VoiceCFM and PitchCFM to generate speech and improve speaker pitch adaptation quality. Experiments show that our method can significantly improve speaker similarity, generating natural and higher-quality speech.
title CycleFlow: Leveraging Cycle Consistency in Flow Matching for Speaker Style Adaptation
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2501.01861