EE-TTS: Emphatic Expressive TTS with Linguistic Information

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhong, Yi, Zhang, Chen, Liu, Xule, Sun, Chenxi, Deng, Weishan, Hu, Haifeng, Sun, Zhongqian
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913856131956736
author Zhong, Yi
Zhang, Chen
Liu, Xule
Sun, Chenxi
Deng, Weishan
Hu, Haifeng
Sun, Zhongqian
author_facet Zhong, Yi
Zhang, Chen
Liu, Xule
Sun, Chenxi
Deng, Weishan
Hu, Haifeng
Sun, Zhongqian
contents While Current TTS systems perform well in synthesizing high-quality speech, producing highly expressive speech remains a challenge. Emphasis, as a critical factor in determining the expressiveness of speech, has attracted more attention nowadays. Previous works usually enhance the emphasis by adding intermediate features, but they can not guarantee the overall expressiveness of the speech. To resolve this matter, we propose Emphatic Expressive TTS (EE-TTS), which leverages multi-level linguistic information from syntax and semantics. EE-TTS contains an emphasis predictor that can identify appropriate emphasis positions from text and a conditioned acoustic model to synthesize expressive speech with emphasis and linguistic information. Experimental results indicate that EE-TTS outperforms baseline with MOS improvements of 0.49 and 0.67 in expressiveness and naturalness. EE-TTS also shows strong generalization across different datasets according to AB test results.
format Preprint
id arxiv_https___arxiv_org_abs_2305_12107
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle EE-TTS: Emphatic Expressive TTS with Linguistic Information
Zhong, Yi
Zhang, Chen
Liu, Xule
Sun, Chenxi
Deng, Weishan
Hu, Haifeng
Sun, Zhongqian
Sound
Computation and Language
Audio and Speech Processing
While Current TTS systems perform well in synthesizing high-quality speech, producing highly expressive speech remains a challenge. Emphasis, as a critical factor in determining the expressiveness of speech, has attracted more attention nowadays. Previous works usually enhance the emphasis by adding intermediate features, but they can not guarantee the overall expressiveness of the speech. To resolve this matter, we propose Emphatic Expressive TTS (EE-TTS), which leverages multi-level linguistic information from syntax and semantics. EE-TTS contains an emphasis predictor that can identify appropriate emphasis positions from text and a conditioned acoustic model to synthesize expressive speech with emphasis and linguistic information. Experimental results indicate that EE-TTS outperforms baseline with MOS improvements of 0.49 and 0.67 in expressiveness and naturalness. EE-TTS also shows strong generalization across different datasets according to AB test results.
title EE-TTS: Emphatic Expressive TTS with Linguistic Information
topic Sound
Computation and Language
Audio and Speech Processing
url https://arxiv.org/abs/2305.12107