Scaling NVIDIA's Multi-speaker Multi-lingual TTS Systems with Zero-Shot TTS to Indic Languages

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Arora, Akshit, Badlani, Rohan, Kim, Sungwon, Valle, Rafael, Catanzaro, Bryan
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910310332366848
author Arora, Akshit
Badlani, Rohan
Kim, Sungwon
Valle, Rafael
Catanzaro, Bryan
author_facet Arora, Akshit
Badlani, Rohan
Kim, Sungwon
Valle, Rafael
Catanzaro, Bryan
contents In this paper, we describe the TTS models developed by NVIDIA for the MMITS-VC (Multi-speaker, Multi-lingual Indic TTS with Voice Cloning) 2024 Challenge. In Tracks 1 and 2, we utilize RAD-MMM to perform few-shot TTS by training additionally on 5 minutes of target speaker data. In Track 3, we utilize P-Flow to perform zero-shot TTS by training on the challenge dataset as well as external datasets. We use HiFi-GAN vocoders for all submissions. RAD-MMM performs competitively on Tracks 1 and 2, while P-Flow ranks first on Track 3, with mean opinion score (MOS) 4.4 and speaker similarity score (SMOS) of 3.62.
format Preprint
id arxiv_https___arxiv_org_abs_2401_13851
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Scaling NVIDIA's Multi-speaker Multi-lingual TTS Systems with Zero-Shot TTS to Indic Languages
Arora, Akshit
Badlani, Rohan
Kim, Sungwon
Valle, Rafael
Catanzaro, Bryan
Sound
Machine Learning
Audio and Speech Processing
In this paper, we describe the TTS models developed by NVIDIA for the MMITS-VC (Multi-speaker, Multi-lingual Indic TTS with Voice Cloning) 2024 Challenge. In Tracks 1 and 2, we utilize RAD-MMM to perform few-shot TTS by training additionally on 5 minutes of target speaker data. In Track 3, we utilize P-Flow to perform zero-shot TTS by training on the challenge dataset as well as external datasets. We use HiFi-GAN vocoders for all submissions. RAD-MMM performs competitively on Tracks 1 and 2, while P-Flow ranks first on Track 3, with mean opinion score (MOS) 4.4 and speaker similarity score (SMOS) of 3.62.
title Scaling NVIDIA's Multi-speaker Multi-lingual TTS Systems with Zero-Shot TTS to Indic Languages
topic Sound
Machine Learning
Audio and Speech Processing
url https://arxiv.org/abs/2401.13851