Accent-VITS:accent transfer for end-to-end TTS
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2023
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866909056706281472 |
|---|---|
| author | Ma, Linhan Zhang, Yongmao Zhu, Xinfa Lei, Yi Ning, Ziqian Zhu, Pengcheng Xie, Lei |
| author_facet | Ma, Linhan Zhang, Yongmao Zhu, Xinfa Lei, Yi Ning, Ziqian Zhu, Pengcheng Xie, Lei |
| contents | Accent transfer aims to transfer an accent from a source speaker to synthetic speech in the target speaker's voice. The main challenge is how to effectively disentangle speaker timbre and accent which are entangled in speech. This paper presents a VITS-based end-to-end accent transfer model named Accent-VITS.Based on the main structure of VITS, Accent-VITS makes substantial improvements to enable effective and stable accent transfer.We leverage a hierarchical CVAE structure to model accent pronunciation information and acoustic features, respectively, using bottleneck features and mel spectrums as constraints.Moreover, the text-to-wave mapping in VITS is decomposed into text-to-accent and accent-to-wave mappings in Accent-VITS. In this way, the disentanglement of accent and speaker timbre becomes be more stable and effective.Experiments on multi-accent and Mandarin datasets show that Accent-VITS achieves higher speaker similarity, accent similarity and speech naturalness as compared with a strong baseline. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2312_16850 |
| institution | arXiv |
| publishDate | 2023 |
| record_format | arxiv |
| spellingShingle | Accent-VITS:accent transfer for end-to-end TTS Ma, Linhan Zhang, Yongmao Zhu, Xinfa Lei, Yi Ning, Ziqian Zhu, Pengcheng Xie, Lei Sound Audio and Speech Processing Accent transfer aims to transfer an accent from a source speaker to synthetic speech in the target speaker's voice. The main challenge is how to effectively disentangle speaker timbre and accent which are entangled in speech. This paper presents a VITS-based end-to-end accent transfer model named Accent-VITS.Based on the main structure of VITS, Accent-VITS makes substantial improvements to enable effective and stable accent transfer.We leverage a hierarchical CVAE structure to model accent pronunciation information and acoustic features, respectively, using bottleneck features and mel spectrums as constraints.Moreover, the text-to-wave mapping in VITS is decomposed into text-to-accent and accent-to-wave mappings in Accent-VITS. In this way, the disentanglement of accent and speaker timbre becomes be more stable and effective.Experiments on multi-accent and Mandarin datasets show that Accent-VITS achieves higher speaker similarity, accent similarity and speech naturalness as compared with a strong baseline. |
| title | Accent-VITS:accent transfer for end-to-end TTS |
| topic | Sound Audio and Speech Processing |
| url | https://arxiv.org/abs/2312.16850 |