A multi-speaker multi-lingual voice cloning system based on vits2 for limmits 2024 challenge

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Wang, Xiaopeng, Lu, Yi, Qi, Xin, Wang, Zhiyong, Xie, Yuankun, Shi, Shuchen, Fu, Ruibo
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866909231562620928
author Wang, Xiaopeng
Lu, Yi
Qi, Xin
Wang, Zhiyong
Xie, Yuankun
Shi, Shuchen
Fu, Ruibo
author_facet Wang, Xiaopeng
Lu, Yi
Qi, Xin
Wang, Zhiyong
Xie, Yuankun
Shi, Shuchen
Fu, Ruibo
contents This paper presents the development of a speech synthesis system for the LIMMITS'24 Challenge, focusing primarily on Track 2. The objective of the challenge is to establish a multi-speaker, multi-lingual Indic Text-to-Speech system with voice cloning capabilities, covering seven Indian languages with both male and female speakers. The system was trained using challenge data and fine-tuned for few-shot voice cloning on target speakers. Evaluation included both mono-lingual and cross-lingual synthesis across all seven languages, with subjective tests assessing naturalness and speaker similarity. Our system uses the VITS2 architecture, augmented with a multi-lingual ID and a BERT model to enhance contextual language comprehension. In Track 1, where no additional data usage was permitted, our model achieved a Speaker Similarity score of 4.02. In Track 2, which allowed the use of extra data, it attained a Speaker Similarity score of 4.17.
format Preprint
id arxiv_https___arxiv_org_abs_2406_17801
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle A multi-speaker multi-lingual voice cloning system based on vits2 for limmits 2024 challenge
Wang, Xiaopeng
Lu, Yi
Qi, Xin
Wang, Zhiyong
Xie, Yuankun
Shi, Shuchen
Fu, Ruibo
Sound
Computation and Language
Audio and Speech Processing
This paper presents the development of a speech synthesis system for the LIMMITS'24 Challenge, focusing primarily on Track 2. The objective of the challenge is to establish a multi-speaker, multi-lingual Indic Text-to-Speech system with voice cloning capabilities, covering seven Indian languages with both male and female speakers. The system was trained using challenge data and fine-tuned for few-shot voice cloning on target speakers. Evaluation included both mono-lingual and cross-lingual synthesis across all seven languages, with subjective tests assessing naturalness and speaker similarity. Our system uses the VITS2 architecture, augmented with a multi-lingual ID and a BERT model to enhance contextual language comprehension. In Track 1, where no additional data usage was permitted, our model achieved a Speaker Similarity score of 4.02. In Track 2, which allowed the use of extra data, it attained a Speaker Similarity score of 4.17.
title A multi-speaker multi-lingual voice cloning system based on vits2 for limmits 2024 challenge
topic Sound
Computation and Language
Audio and Speech Processing
url https://arxiv.org/abs/2406.17801