Word-Level Emotional Expression Control in Zero-Shot Text-to-Speech Synthesis

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Tianrui, Wang, Haoyu, Ge, Meng, Gong, Cheng, Qiang, Chunyu, Ma, Ziyang, Huang, Zikang, Yang, Guanrou, Wang, Xiaobao, Chng, Eng Siong, Chen, Xie, Wang, Longbiao, Dang, Jianwu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915719546929152
author Wang, Tianrui
Wang, Haoyu
Ge, Meng
Gong, Cheng
Qiang, Chunyu
Ma, Ziyang
Huang, Zikang
Yang, Guanrou
Wang, Xiaobao
Chng, Eng Siong
Chen, Xie
Wang, Longbiao
Dang, Jianwu
author_facet Wang, Tianrui
Wang, Haoyu
Ge, Meng
Gong, Cheng
Qiang, Chunyu
Ma, Ziyang
Huang, Zikang
Yang, Guanrou
Wang, Xiaobao
Chng, Eng Siong
Chen, Xie
Wang, Longbiao
Dang, Jianwu
contents While emotional text-to-speech (TTS) has made significant progress, most existing research remains limited to utterance-level emotional expression and fails to support word-level control. Achieving word-level expressive control poses fundamental challenges, primarily due to the complexity of modeling multi-emotion transitions and the scarcity of annotated datasets that capture intra-sentence emotional and prosodic variation. In this paper, we propose WeSCon, the first self-training framework that enables word-level control of both emotion and speaking rate in a pretrained zero-shot TTS model, without relying on datasets containing intra-sentence emotion or speed transitions. Our method introduces a transition-smoothing strategy and a dynamic speed control mechanism to guide the pretrained TTS model in performing word-level expressive synthesis through a multi-round inference process. To further simplify the inference, we incorporate a dynamic emotional attention bias mechanism and fine-tune the model via self-training, thereby activating its ability for word-level expressive control in an end-to-end manner. Experimental results show that WeSCon effectively overcomes data scarcity, achieving state-of-the-art performance in word-level emotional expression control while preserving the strong zero-shot synthesis capabilities of the original TTS model.
format Preprint
id arxiv_https___arxiv_org_abs_2509_24629
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Word-Level Emotional Expression Control in Zero-Shot Text-to-Speech Synthesis
Wang, Tianrui
Wang, Haoyu
Ge, Meng
Gong, Cheng
Qiang, Chunyu
Ma, Ziyang
Huang, Zikang
Yang, Guanrou
Wang, Xiaobao
Chng, Eng Siong
Chen, Xie
Wang, Longbiao
Dang, Jianwu
Audio and Speech Processing
Sound
While emotional text-to-speech (TTS) has made significant progress, most existing research remains limited to utterance-level emotional expression and fails to support word-level control. Achieving word-level expressive control poses fundamental challenges, primarily due to the complexity of modeling multi-emotion transitions and the scarcity of annotated datasets that capture intra-sentence emotional and prosodic variation. In this paper, we propose WeSCon, the first self-training framework that enables word-level control of both emotion and speaking rate in a pretrained zero-shot TTS model, without relying on datasets containing intra-sentence emotion or speed transitions. Our method introduces a transition-smoothing strategy and a dynamic speed control mechanism to guide the pretrained TTS model in performing word-level expressive synthesis through a multi-round inference process. To further simplify the inference, we incorporate a dynamic emotional attention bias mechanism and fine-tune the model via self-training, thereby activating its ability for word-level expressive control in an end-to-end manner. Experimental results show that WeSCon effectively overcomes data scarcity, achieving state-of-the-art performance in word-level emotional expression control while preserving the strong zero-shot synthesis capabilities of the original TTS model.
title Word-Level Emotional Expression Control in Zero-Shot Text-to-Speech Synthesis
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2509.24629