Multi-Modal Automatic Prosody Annotation with Contrastive Pretraining of SSWP

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhong, Jinzuomu, Li, Yang, Huang, Hui, Richmond, Korin, Liu, Jie, Su, Zhiba, Guo, Jing, Tang, Benlai, Zhu, Fengjie
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909221374656512
author Zhong, Jinzuomu
Li, Yang
Huang, Hui
Richmond, Korin
Liu, Jie
Su, Zhiba
Guo, Jing
Tang, Benlai
Zhu, Fengjie
author_facet Zhong, Jinzuomu
Li, Yang
Huang, Hui
Richmond, Korin
Liu, Jie
Su, Zhiba
Guo, Jing
Tang, Benlai
Zhu, Fengjie
contents In expressive and controllable Text-to-Speech (TTS), explicit prosodic features significantly improve the naturalness and controllability of synthesised speech. However, manual prosody annotation is labor-intensive and inconsistent. To address this issue, a two-stage automatic annotation pipeline is novelly proposed in this paper. In the first stage, we use contrastive pretraining of Speech-Silence and Word-Punctuation (SSWP) pairs to enhance prosodic information in latent representations. In the second stage, we build a multi-modal prosody annotator, comprising pretrained encoders, a text-speech fusing scheme, and a sequence classifier. Experiments on English prosodic boundaries demonstrate that our method achieves state-of-the-art (SOTA) performance with 0.72 and 0.93 f1 score for Prosodic Word and Prosodic Phrase boundary respectively, while bearing remarkable robustness to data scarcity.
format Preprint
id arxiv_https___arxiv_org_abs_2309_05423
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Multi-Modal Automatic Prosody Annotation with Contrastive Pretraining of SSWP
Zhong, Jinzuomu
Li, Yang
Huang, Hui
Richmond, Korin
Liu, Jie
Su, Zhiba
Guo, Jing
Tang, Benlai
Zhu, Fengjie
Audio and Speech Processing
Artificial Intelligence
Computation and Language
Sound
In expressive and controllable Text-to-Speech (TTS), explicit prosodic features significantly improve the naturalness and controllability of synthesised speech. However, manual prosody annotation is labor-intensive and inconsistent. To address this issue, a two-stage automatic annotation pipeline is novelly proposed in this paper. In the first stage, we use contrastive pretraining of Speech-Silence and Word-Punctuation (SSWP) pairs to enhance prosodic information in latent representations. In the second stage, we build a multi-modal prosody annotator, comprising pretrained encoders, a text-speech fusing scheme, and a sequence classifier. Experiments on English prosodic boundaries demonstrate that our method achieves state-of-the-art (SOTA) performance with 0.72 and 0.93 f1 score for Prosodic Word and Prosodic Phrase boundary respectively, while bearing remarkable robustness to data scarcity.
title Multi-Modal Automatic Prosody Annotation with Contrastive Pretraining of SSWP
topic Audio and Speech Processing
Artificial Intelligence
Computation and Language
Sound
url https://arxiv.org/abs/2309.05423