Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yoon, Jinsung, Jeong, Wooyeol, Gim, Jio, Suh, Young-Joo
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908484279205888
author Yoon, Jinsung
Jeong, Wooyeol
Gim, Jio
Suh, Young-Joo
author_facet Yoon, Jinsung
Jeong, Wooyeol
Gim, Jio
Suh, Young-Joo
contents Emotional voice conversion (EVC) aims to modify the emotional style of speech while preserving its linguistic content. In practical EVC, controllability, the ability to independently control speaker identity and emotional style using distinct references, is crucial. However, existing methods often struggle to fully disentangle these attributes and lack the ability to model fine-grained emotional expressions such as temporal dynamics. We propose Maestro-EVC, a controllable EVC framework that enables independent control of content, speaker identity, and emotion by effectively disentangling each attribute from separate references. We further introduce a temporal emotion representation and an explicit prosody modeling with prosody augmentation to robustly capture and transfer the temporal dynamics of the target emotion, even under prosody-mismatched conditions. Experimental results confirm that Maestro-EVC achieves high-quality, controllable, and emotionally expressive speech synthesis.
format Preprint
id arxiv_https___arxiv_org_abs_2508_06890
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody
Yoon, Jinsung
Jeong, Wooyeol
Gim, Jio
Suh, Young-Joo
Sound
Artificial Intelligence
Computation and Language
Audio and Speech Processing
Emotional voice conversion (EVC) aims to modify the emotional style of speech while preserving its linguistic content. In practical EVC, controllability, the ability to independently control speaker identity and emotional style using distinct references, is crucial. However, existing methods often struggle to fully disentangle these attributes and lack the ability to model fine-grained emotional expressions such as temporal dynamics. We propose Maestro-EVC, a controllable EVC framework that enables independent control of content, speaker identity, and emotion by effectively disentangling each attribute from separate references. We further introduce a temporal emotion representation and an explicit prosody modeling with prosody augmentation to robustly capture and transfer the temporal dynamics of the target emotion, even under prosody-mismatched conditions. Experimental results confirm that Maestro-EVC achieves high-quality, controllable, and emotionally expressive speech synthesis.
title Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody
topic Sound
Artificial Intelligence
Computation and Language
Audio and Speech Processing
url https://arxiv.org/abs/2508.06890