ZSDEVC: Zero-Shot Diffusion-based Emotional Voice Conversion with Disentangled Mechanism

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Chou, Hsing-Hang, Lin, Yun-Shao, Sung, Ching-Chin, Tsao, Yu, Lee, Chi-Chun
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866916970520117248
author Chou, Hsing-Hang
Lin, Yun-Shao
Sung, Ching-Chin
Tsao, Yu
Lee, Chi-Chun
author_facet Chou, Hsing-Hang
Lin, Yun-Shao
Sung, Ching-Chin
Tsao, Yu
Lee, Chi-Chun
contents The human voice conveys not just words but also emotional states and individuality. Emotional voice conversion (EVC) modifies emotional expressions while preserving linguistic content and speaker identity, improving applications like human-machine interaction. While deep learning has advanced EVC models for specific target speakers on well-crafted emotional datasets, existing methods often face issues with emotion accuracy and speech distortion. In addition, the zero-shot scenario, in which emotion conversion is applied to unseen speakers, remains underexplored. This work introduces a novel diffusion framework with disentangled mechanisms and expressive guidance, trained on a large emotional speech dataset and evaluated on unseen speakers across in-domain and out-of-domain datasets. Experimental results show that our method produces expressive speech with high emotional accuracy, naturalness, and quality, showcasing its potential for broader EVC applications.
format Preprint
id arxiv_https___arxiv_org_abs_2409_03636
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle ZSDEVC: Zero-Shot Diffusion-based Emotional Voice Conversion with Disentangled Mechanism
Chou, Hsing-Hang
Lin, Yun-Shao
Sung, Ching-Chin
Tsao, Yu
Lee, Chi-Chun
Audio and Speech Processing
The human voice conveys not just words but also emotional states and individuality. Emotional voice conversion (EVC) modifies emotional expressions while preserving linguistic content and speaker identity, improving applications like human-machine interaction. While deep learning has advanced EVC models for specific target speakers on well-crafted emotional datasets, existing methods often face issues with emotion accuracy and speech distortion. In addition, the zero-shot scenario, in which emotion conversion is applied to unseen speakers, remains underexplored. This work introduces a novel diffusion framework with disentangled mechanisms and expressive guidance, trained on a large emotional speech dataset and evaluated on unseen speakers across in-domain and out-of-domain datasets. Experimental results show that our method produces expressive speech with high emotional accuracy, naturalness, and quality, showcasing its potential for broader EVC applications.
title ZSDEVC: Zero-Shot Diffusion-based Emotional Voice Conversion with Disentangled Mechanism
topic Audio and Speech Processing
url https://arxiv.org/abs/2409.03636