Rasa: Building Expressive Speech Synthesis Systems for Indian Languages in Low-resource Settings

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Varadhan, Praveen Srinivasa, Sankar, Ashwin, Raju, Giri, Khapra, Mitesh M.
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929480744828928
author Varadhan, Praveen Srinivasa
Sankar, Ashwin
Raju, Giri
Khapra, Mitesh M.
author_facet Varadhan, Praveen Srinivasa
Sankar, Ashwin
Raju, Giri
Khapra, Mitesh M.
contents We release Rasa, the first multilingual expressive TTS dataset for any Indian language, which contains 10 hours of neutral speech and 1-3 hours of expressive speech for each of the 6 Ekman emotions covering 3 languages: Assamese, Bengali, & Tamil. Our ablation studies reveal that just 1 hour of neutral and 30 minutes of expressive data can yield a Fair system as indicated by MUSHRA scores. Increasing neutral data to 10 hours, with minimal expressive data, significantly enhances expressiveness. This offers a practical recipe for resource-constrained languages, prioritizing easily obtainable neutral data alongside smaller amounts of expressive data. We show the importance of syllabically balanced data and pooling emotions to enhance expressiveness. We also highlight challenges in generating specific emotions, e.g., fear and surprise.
format Preprint
id arxiv_https___arxiv_org_abs_2407_14056
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Rasa: Building Expressive Speech Synthesis Systems for Indian Languages in Low-resource Settings
Varadhan, Praveen Srinivasa
Sankar, Ashwin
Raju, Giri
Khapra, Mitesh M.
Computation and Language
Machine Learning
Sound
Audio and Speech Processing
We release Rasa, the first multilingual expressive TTS dataset for any Indian language, which contains 10 hours of neutral speech and 1-3 hours of expressive speech for each of the 6 Ekman emotions covering 3 languages: Assamese, Bengali, & Tamil. Our ablation studies reveal that just 1 hour of neutral and 30 minutes of expressive data can yield a Fair system as indicated by MUSHRA scores. Increasing neutral data to 10 hours, with minimal expressive data, significantly enhances expressiveness. This offers a practical recipe for resource-constrained languages, prioritizing easily obtainable neutral data alongside smaller amounts of expressive data. We show the importance of syllabically balanced data and pooling emotions to enhance expressiveness. We also highlight challenges in generating specific emotions, e.g., fear and surprise.
title Rasa: Building Expressive Speech Synthesis Systems for Indian Languages in Low-resource Settings
topic Computation and Language
Machine Learning
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2407.14056