CoCoEmo: Composable and Controllable Human-Like Emotional TTS via Activation Steering

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Wang, Siyi, Tan, Shihong, Liu, Siyi, Jia, Hong, Huang, Gongping, Bailey, James, Dang, Ting
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908808242003968
author Wang, Siyi
Tan, Shihong
Liu, Siyi
Jia, Hong
Huang, Gongping
Bailey, James
Dang, Ting
author_facet Wang, Siyi
Tan, Shihong
Liu, Siyi
Jia, Hong
Huang, Gongping
Bailey, James
Dang, Ting
contents Emotional expression in human speech is nuanced and compositional, often involving multiple, sometimes conflicting, affective cues that may diverge from linguistic content. In contrast, most expressive text-to-speech systems enforce a single utterance-level emotion, collapsing affective diversity and suppressing mixed or text-emotion-misaligned expression. While activation steering via latent direction vectors offers a promising solution, it remains unclear whether emotion representations are linearly steerable in TTS, where steering should be applied within hybrid TTS architectures, and how such complex emotion behaviors should be evaluated. This paper presents the first systematic analysis of activation steering for emotional control in hybrid TTS models, introducing a quantitative, controllable steering framework, and multi-rater evaluation protocols that enable composable mixed-emotion synthesis and reliable text-emotion mismatch synthesis. Our results demonstrate, for the first time, that emotional prosody and expressive variability are primarily synthesized by the TTS language module instead of the flow-matching module, and also provide a lightweight steering approach for generating natural, human-like emotional speech.
format Preprint
id arxiv_https___arxiv_org_abs_2602_03420
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle CoCoEmo: Composable and Controllable Human-Like Emotional TTS via Activation Steering
Wang, Siyi
Tan, Shihong
Liu, Siyi
Jia, Hong
Huang, Gongping
Bailey, James
Dang, Ting
Sound
Machine Learning
Emotional expression in human speech is nuanced and compositional, often involving multiple, sometimes conflicting, affective cues that may diverge from linguistic content. In contrast, most expressive text-to-speech systems enforce a single utterance-level emotion, collapsing affective diversity and suppressing mixed or text-emotion-misaligned expression. While activation steering via latent direction vectors offers a promising solution, it remains unclear whether emotion representations are linearly steerable in TTS, where steering should be applied within hybrid TTS architectures, and how such complex emotion behaviors should be evaluated. This paper presents the first systematic analysis of activation steering for emotional control in hybrid TTS models, introducing a quantitative, controllable steering framework, and multi-rater evaluation protocols that enable composable mixed-emotion synthesis and reliable text-emotion mismatch synthesis. Our results demonstrate, for the first time, that emotional prosody and expressive variability are primarily synthesized by the TTS language module instead of the flow-matching module, and also provide a lightweight steering approach for generating natural, human-like emotional speech.
title CoCoEmo: Composable and Controllable Human-Like Emotional TTS via Activation Steering
topic Sound
Machine Learning
url https://arxiv.org/abs/2602.03420