ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Pan, Yu, Hu, Yanni, Yang, Yuguang, Yao, Jixun, Ye, Jianhao, Zhou, Hongbin, Ma, Lei, Zhao, Jianjun
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910955472945152
author Pan, Yu
Hu, Yanni
Yang, Yuguang
Yao, Jixun
Ye, Jianhao
Zhou, Hongbin
Ma, Lei
Zhao, Jianjun
author_facet Pan, Yu
Hu, Yanni
Yang, Yuguang
Yao, Jixun
Ye, Jianhao
Zhou, Hongbin
Ma, Lei
Zhao, Jianjun
contents Despite great advances, achieving high-fidelity emotional voice conversion (EVC) with flexible and interpretable control remains challenging. This paper introduces ClapFM-EVC, a novel EVC framework capable of generating high-quality converted speech driven by natural language prompts or reference speech with adjustable emotion intensity. We first propose EVC-CLAP, an emotional contrastive language-audio pre-training model, guided by natural language prompts and categorical labels, to extract and align fine-grained emotional elements across speech and text modalities. Then, a FuEncoder with an adaptive intensity gate is presented to seamless fuse emotional features with Phonetic PosteriorGrams from a pre-trained ASR model. To further improve emotion expressiveness and speech naturalness, we propose a flow matching model conditioned on these captured features to reconstruct Mel-spectrogram of source speech. Subjective and objective evaluations validate the effectiveness of ClapFM-EVC.
format Preprint
id arxiv_https___arxiv_org_abs_2505_13805
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech
Pan, Yu
Hu, Yanni
Yang, Yuguang
Yao, Jixun
Ye, Jianhao
Zhou, Hongbin
Ma, Lei
Zhao, Jianjun
Sound
Artificial Intelligence
Audio and Speech Processing
Despite great advances, achieving high-fidelity emotional voice conversion (EVC) with flexible and interpretable control remains challenging. This paper introduces ClapFM-EVC, a novel EVC framework capable of generating high-quality converted speech driven by natural language prompts or reference speech with adjustable emotion intensity. We first propose EVC-CLAP, an emotional contrastive language-audio pre-training model, guided by natural language prompts and categorical labels, to extract and align fine-grained emotional elements across speech and text modalities. Then, a FuEncoder with an adaptive intensity gate is presented to seamless fuse emotional features with Phonetic PosteriorGrams from a pre-trained ASR model. To further improve emotion expressiveness and speech naturalness, we propose a flow matching model conditioned on these captured features to reconstruct Mel-spectrogram of source speech. Subjective and objective evaluations validate the effectiveness of ClapFM-EVC.
title ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech
topic Sound
Artificial Intelligence
Audio and Speech Processing
url https://arxiv.org/abs/2505.13805