TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Jeong, Jaeseok, Lee, Yuna, Kwon, Mingi, Uh, Youngjung
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866909677655162880
author Jeong, Jaeseok
Lee, Yuna
Kwon, Mingi
Uh, Youngjung
author_facet Jeong, Jaeseok
Lee, Yuna
Kwon, Mingi
Uh, Youngjung
contents Recent advances in text-to-speech (TTS) have enabled natural speech synthesis, but fine-grained, time-varying emotion control remains challenging. Existing methods often allow only utterance-level control and require full model fine-tuning with a large emotion speech dataset, which can degrade performance. Inspired by adding conditional control to the existing model in ControlNet (Zhang et al, 2023), we propose the first ControlNet-based approach for controllable flow-matching TTS (TTS-CtrlNet), which freezes the original model and introduces a trainable copy of it to process additional conditions. We show that TTS-CtrlNet can boost the pretrained large TTS model by adding intuitive, scalable, and time-varying emotion control while inheriting the ability of the original model (e.g., zero-shot voice cloning & naturalness). Furthermore, we provide practical recipes for adding emotion control: 1) optimal architecture design choice with block analysis, 2) emotion-specific flow step, and 3) flexible control scale. Experiments show that ours can effectively add an emotion controller to existing TTS, and achieves state-of-the-art performance with emotion similarity scores: Emo-SIM and Aro-Val SIM. The project page is available at: https://curryjung.github.io/ttsctrlnet_project_page
format Preprint
id arxiv_https___arxiv_org_abs_2507_04349
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet
Jeong, Jaeseok
Lee, Yuna
Kwon, Mingi
Uh, Youngjung
Sound
Audio and Speech Processing
Recent advances in text-to-speech (TTS) have enabled natural speech synthesis, but fine-grained, time-varying emotion control remains challenging. Existing methods often allow only utterance-level control and require full model fine-tuning with a large emotion speech dataset, which can degrade performance. Inspired by adding conditional control to the existing model in ControlNet (Zhang et al, 2023), we propose the first ControlNet-based approach for controllable flow-matching TTS (TTS-CtrlNet), which freezes the original model and introduces a trainable copy of it to process additional conditions. We show that TTS-CtrlNet can boost the pretrained large TTS model by adding intuitive, scalable, and time-varying emotion control while inheriting the ability of the original model (e.g., zero-shot voice cloning & naturalness). Furthermore, we provide practical recipes for adding emotion control: 1) optimal architecture design choice with block analysis, 2) emotion-specific flow step, and 3) flexible control scale. Experiments show that ours can effectively add an emotion controller to existing TTS, and achieves state-of-the-art performance with emotion similarity scores: Emo-SIM and Aro-Val SIM. The project page is available at: https://curryjung.github.io/ttsctrlnet_project_page
title TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2507.04349