Accent-VITS:accent transfer for end-to-end TTS

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Ma, Linhan, Zhang, Yongmao, Zhu, Xinfa, Lei, Yi, Ning, Ziqian, Zhu, Pengcheng, Xie, Lei
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909056706281472
author Ma, Linhan
Zhang, Yongmao
Zhu, Xinfa
Lei, Yi
Ning, Ziqian
Zhu, Pengcheng
Xie, Lei
author_facet Ma, Linhan
Zhang, Yongmao
Zhu, Xinfa
Lei, Yi
Ning, Ziqian
Zhu, Pengcheng
Xie, Lei
contents Accent transfer aims to transfer an accent from a source speaker to synthetic speech in the target speaker's voice. The main challenge is how to effectively disentangle speaker timbre and accent which are entangled in speech. This paper presents a VITS-based end-to-end accent transfer model named Accent-VITS.Based on the main structure of VITS, Accent-VITS makes substantial improvements to enable effective and stable accent transfer.We leverage a hierarchical CVAE structure to model accent pronunciation information and acoustic features, respectively, using bottleneck features and mel spectrums as constraints.Moreover, the text-to-wave mapping in VITS is decomposed into text-to-accent and accent-to-wave mappings in Accent-VITS. In this way, the disentanglement of accent and speaker timbre becomes be more stable and effective.Experiments on multi-accent and Mandarin datasets show that Accent-VITS achieves higher speaker similarity, accent similarity and speech naturalness as compared with a strong baseline.
format Preprint
id arxiv_https___arxiv_org_abs_2312_16850
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Accent-VITS:accent transfer for end-to-end TTS
Ma, Linhan
Zhang, Yongmao
Zhu, Xinfa
Lei, Yi
Ning, Ziqian
Zhu, Pengcheng
Xie, Lei
Sound
Audio and Speech Processing
Accent transfer aims to transfer an accent from a source speaker to synthetic speech in the target speaker's voice. The main challenge is how to effectively disentangle speaker timbre and accent which are entangled in speech. This paper presents a VITS-based end-to-end accent transfer model named Accent-VITS.Based on the main structure of VITS, Accent-VITS makes substantial improvements to enable effective and stable accent transfer.We leverage a hierarchical CVAE structure to model accent pronunciation information and acoustic features, respectively, using bottleneck features and mel spectrums as constraints.Moreover, the text-to-wave mapping in VITS is decomposed into text-to-accent and accent-to-wave mappings in Accent-VITS. In this way, the disentanglement of accent and speaker timbre becomes be more stable and effective.Experiments on multi-accent and Mandarin datasets show that Accent-VITS achieves higher speaker similarity, accent similarity and speech naturalness as compared with a strong baseline.
title Accent-VITS:accent transfer for end-to-end TTS
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2312.16850