EmoKnob: Enhance Voice Cloning with Fine-Grained Emotion Control

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Haozhe, Chen, Run, Hirschberg, Julia
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916417384742912
author Chen, Haozhe
Chen, Run
Hirschberg, Julia
author_facet Chen, Haozhe
Chen, Run
Hirschberg, Julia
contents While recent advances in Text-to-Speech (TTS) technology produce natural and expressive speech, they lack the option for users to select emotion and control intensity. We propose EmoKnob, a framework that allows fine-grained emotion control in speech synthesis with few-shot demonstrative samples of arbitrary emotion. Our framework leverages the expressive speaker representation space made possible by recent advances in foundation voice cloning models. Based on the few-shot capability of our emotion control framework, we propose two methods to apply emotion control on emotions described by open-ended text, enabling an intuitive interface for controlling a diverse array of nuanced emotions. To facilitate a more systematic emotional speech synthesis field, we introduce a set of evaluation metrics designed to rigorously assess the faithfulness and recognizability of emotion control frameworks. Through objective and subjective evaluations, we show that our emotion control framework effectively embeds emotions into speech and surpasses emotion expressiveness of commercial TTS services.
format Preprint
id arxiv_https___arxiv_org_abs_2410_00316
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle EmoKnob: Enhance Voice Cloning with Fine-Grained Emotion Control
Chen, Haozhe
Chen, Run
Hirschberg, Julia
Computation and Language
Artificial Intelligence
Human-Computer Interaction
Sound
Audio and Speech Processing
While recent advances in Text-to-Speech (TTS) technology produce natural and expressive speech, they lack the option for users to select emotion and control intensity. We propose EmoKnob, a framework that allows fine-grained emotion control in speech synthesis with few-shot demonstrative samples of arbitrary emotion. Our framework leverages the expressive speaker representation space made possible by recent advances in foundation voice cloning models. Based on the few-shot capability of our emotion control framework, we propose two methods to apply emotion control on emotions described by open-ended text, enabling an intuitive interface for controlling a diverse array of nuanced emotions. To facilitate a more systematic emotional speech synthesis field, we introduce a set of evaluation metrics designed to rigorously assess the faithfulness and recognizability of emotion control frameworks. Through objective and subjective evaluations, we show that our emotion control framework effectively embeds emotions into speech and surpasses emotion expressiveness of commercial TTS services.
title EmoKnob: Enhance Voice Cloning with Fine-Grained Emotion Control
topic Computation and Language
Artificial Intelligence
Human-Computer Interaction
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2410.00316