VECL-TTS: Voice identity and Emotional style controllable Cross-Lingual Text-to-Speech

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Gudmalwar, Ashishkumar, Shah, Nirmesh, Akarsh, Sai, Wasnik, Pankaj, Shah, Rajiv Ratn
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916283629436928
author Gudmalwar, Ashishkumar
Shah, Nirmesh
Akarsh, Sai
Wasnik, Pankaj
Shah, Rajiv Ratn
author_facet Gudmalwar, Ashishkumar
Shah, Nirmesh
Akarsh, Sai
Wasnik, Pankaj
Shah, Rajiv Ratn
contents Despite the significant advancements in Text-to-Speech (TTS) systems, their full utilization in automatic dubbing remains limited. This task necessitates the extraction of voice identity and emotional style from a reference speech in a source language and subsequently transferring them to a target language using cross-lingual TTS techniques. While previous approaches have mainly concentrated on controlling voice identity within the cross-lingual TTS framework, there has been limited work on incorporating emotion and voice identity together. To this end, we introduce an end-to-end Voice Identity and Emotional Style Controllable Cross-Lingual (VECL) TTS system using multilingual speakers and an emotion embedding network. Moreover, we introduce content and style consistency losses to enhance the quality of synthesized speech further. The proposed system achieved an average relative improvement of 8.83\% compared to the state-of-the-art (SOTA) methods on a database comprising English and three Indian languages (Hindi, Telugu, and Marathi).
format Preprint
id arxiv_https___arxiv_org_abs_2406_08076
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle VECL-TTS: Voice identity and Emotional style controllable Cross-Lingual Text-to-Speech
Gudmalwar, Ashishkumar
Shah, Nirmesh
Akarsh, Sai
Wasnik, Pankaj
Shah, Rajiv Ratn
Audio and Speech Processing
Sound
Despite the significant advancements in Text-to-Speech (TTS) systems, their full utilization in automatic dubbing remains limited. This task necessitates the extraction of voice identity and emotional style from a reference speech in a source language and subsequently transferring them to a target language using cross-lingual TTS techniques. While previous approaches have mainly concentrated on controlling voice identity within the cross-lingual TTS framework, there has been limited work on incorporating emotion and voice identity together. To this end, we introduce an end-to-end Voice Identity and Emotional Style Controllable Cross-Lingual (VECL) TTS system using multilingual speakers and an emotion embedding network. Moreover, we introduce content and style consistency losses to enhance the quality of synthesized speech further. The proposed system achieved an average relative improvement of 8.83\% compared to the state-of-the-art (SOTA) methods on a database comprising English and three Indian languages (Hindi, Telugu, and Marathi).
title VECL-TTS: Voice identity and Emotional style controllable Cross-Lingual Text-to-Speech
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2406.08076