UMETTS: A Unified Framework for Emotional Text-to-Speech Synthesis with Multimodal Prompts

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Cheng, Zhi-Qi, Li, Xiang, He, Jun-Yan, Chen, Junyao, Fan, Xiaomao, Peng, Xiaojiang, Hauptmann, Alexander G.
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912236715376640
author Cheng, Zhi-Qi
Li, Xiang
He, Jun-Yan
Chen, Junyao
Fan, Xiaomao
Peng, Xiaojiang
Hauptmann, Alexander G.
author_facet Cheng, Zhi-Qi
Li, Xiang
He, Jun-Yan
Chen, Junyao
Fan, Xiaomao
Peng, Xiaojiang
Hauptmann, Alexander G.
contents Emotional Text-to-Speech (E-TTS) synthesis has garnered significant attention in recent years due to its potential to revolutionize human-computer interaction. However, current E-TTS approaches often struggle to capture the intricacies of human emotions, primarily relying on oversimplified emotional labels or single-modality input. In this paper, we introduce the Unified Multimodal Prompt-Induced Emotional Text-to-Speech System (UMETTS), a novel framework that leverages emotional cues from multiple modalities to generate highly expressive and emotionally resonant speech. The core of UMETTS consists of two key components: the Emotion Prompt Alignment Module (EP-Align) and the Emotion Embedding-Induced TTS Module (EMI-TTS). (1) EP-Align employs contrastive learning to align emotional features across text, audio, and visual modalities, ensuring a coherent fusion of multimodal information. (2) Subsequently, EMI-TTS integrates the aligned emotional embeddings with state-of-the-art TTS models to synthesize speech that accurately reflects the intended emotions. Extensive evaluations show that UMETTS achieves significant improvements in emotion accuracy and speech naturalness, outperforming traditional E-TTS methods on both objective and subjective metrics.
format Preprint
id arxiv_https___arxiv_org_abs_2404_18398
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle UMETTS: A Unified Framework for Emotional Text-to-Speech Synthesis with Multimodal Prompts
Cheng, Zhi-Qi
Li, Xiang
He, Jun-Yan
Chen, Junyao
Fan, Xiaomao
Peng, Xiaojiang
Hauptmann, Alexander G.
Computation and Language
Multimedia
Emotional Text-to-Speech (E-TTS) synthesis has garnered significant attention in recent years due to its potential to revolutionize human-computer interaction. However, current E-TTS approaches often struggle to capture the intricacies of human emotions, primarily relying on oversimplified emotional labels or single-modality input. In this paper, we introduce the Unified Multimodal Prompt-Induced Emotional Text-to-Speech System (UMETTS), a novel framework that leverages emotional cues from multiple modalities to generate highly expressive and emotionally resonant speech. The core of UMETTS consists of two key components: the Emotion Prompt Alignment Module (EP-Align) and the Emotion Embedding-Induced TTS Module (EMI-TTS). (1) EP-Align employs contrastive learning to align emotional features across text, audio, and visual modalities, ensuring a coherent fusion of multimodal information. (2) Subsequently, EMI-TTS integrates the aligned emotional embeddings with state-of-the-art TTS models to synthesize speech that accurately reflects the intended emotions. Extensive evaluations show that UMETTS achieves significant improvements in emotion accuracy and speech naturalness, outperforming traditional E-TTS methods on both objective and subjective metrics.
title UMETTS: A Unified Framework for Emotional Text-to-Speech Synthesis with Multimodal Prompts
topic Computation and Language
Multimedia
url https://arxiv.org/abs/2404.18398