UMETTS: A Unified Framework for Emotional Text-to-Speech Synthesis with Multimodal Prompts
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866912236715376640 |
|---|---|
| author | Cheng, Zhi-Qi Li, Xiang He, Jun-Yan Chen, Junyao Fan, Xiaomao Peng, Xiaojiang Hauptmann, Alexander G. |
| author_facet | Cheng, Zhi-Qi Li, Xiang He, Jun-Yan Chen, Junyao Fan, Xiaomao Peng, Xiaojiang Hauptmann, Alexander G. |
| contents | Emotional Text-to-Speech (E-TTS) synthesis has garnered significant attention in recent years due to its potential to revolutionize human-computer interaction. However, current E-TTS approaches often struggle to capture the intricacies of human emotions, primarily relying on oversimplified emotional labels or single-modality input. In this paper, we introduce the Unified Multimodal Prompt-Induced Emotional Text-to-Speech System (UMETTS), a novel framework that leverages emotional cues from multiple modalities to generate highly expressive and emotionally resonant speech. The core of UMETTS consists of two key components: the Emotion Prompt Alignment Module (EP-Align) and the Emotion Embedding-Induced TTS Module (EMI-TTS). (1) EP-Align employs contrastive learning to align emotional features across text, audio, and visual modalities, ensuring a coherent fusion of multimodal information. (2) Subsequently, EMI-TTS integrates the aligned emotional embeddings with state-of-the-art TTS models to synthesize speech that accurately reflects the intended emotions. Extensive evaluations show that UMETTS achieves significant improvements in emotion accuracy and speech naturalness, outperforming traditional E-TTS methods on both objective and subjective metrics. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2404_18398 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | UMETTS: A Unified Framework for Emotional Text-to-Speech Synthesis with Multimodal Prompts Cheng, Zhi-Qi Li, Xiang He, Jun-Yan Chen, Junyao Fan, Xiaomao Peng, Xiaojiang Hauptmann, Alexander G. Computation and Language Multimedia Emotional Text-to-Speech (E-TTS) synthesis has garnered significant attention in recent years due to its potential to revolutionize human-computer interaction. However, current E-TTS approaches often struggle to capture the intricacies of human emotions, primarily relying on oversimplified emotional labels or single-modality input. In this paper, we introduce the Unified Multimodal Prompt-Induced Emotional Text-to-Speech System (UMETTS), a novel framework that leverages emotional cues from multiple modalities to generate highly expressive and emotionally resonant speech. The core of UMETTS consists of two key components: the Emotion Prompt Alignment Module (EP-Align) and the Emotion Embedding-Induced TTS Module (EMI-TTS). (1) EP-Align employs contrastive learning to align emotional features across text, audio, and visual modalities, ensuring a coherent fusion of multimodal information. (2) Subsequently, EMI-TTS integrates the aligned emotional embeddings with state-of-the-art TTS models to synthesize speech that accurately reflects the intended emotions. Extensive evaluations show that UMETTS achieves significant improvements in emotion accuracy and speech naturalness, outperforming traditional E-TTS methods on both objective and subjective metrics. |
| title | UMETTS: A Unified Framework for Emotional Text-to-Speech Synthesis with Multimodal Prompts |
| topic | Computation and Language Multimedia |
| url | https://arxiv.org/abs/2404.18398 |