Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Wang, Yongqi, Zhang, Chunlei, Chen, Hangting, Zhao, Zhou, Yu, Dong
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866918044503113728
author Wang, Yongqi
Zhang, Chunlei
Chen, Hangting
Zhao, Zhou
Yu, Dong
author_facet Wang, Yongqi
Zhang, Chunlei
Chen, Hangting
Zhao, Zhou
Yu, Dong
contents Controllable TTS models with natural language prompts often lack the ability for fine-grained control and face a scarcity of high-quality data. We propose a two-stage style-controllable TTS system with language models, utilizing a quantized masked-autoencoded style-rich representation as an intermediary. In the first stage, an autoregressive transformer is used for the conditional generation of these style-rich tokens from text and control signals. The second stage generates codec tokens from both text and sampled style-rich tokens. Experiments show that training the first-stage model on extensive datasets enhances the content robustness of the two-stage model as well as control capabilities over multiple attributes. By selectively combining discrete labels and speaker embeddings, we explore fully controlling the speaker's timbre and other stylistic information, and adjusting attributes like emotion for a specified speaker. Audio samples are available at https://style-ar-tts.github.io.
format Preprint
id arxiv_https___arxiv_org_abs_2506_02997
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation
Wang, Yongqi
Zhang, Chunlei
Chen, Hangting
Zhao, Zhou
Yu, Dong
Multimedia
Controllable TTS models with natural language prompts often lack the ability for fine-grained control and face a scarcity of high-quality data. We propose a two-stage style-controllable TTS system with language models, utilizing a quantized masked-autoencoded style-rich representation as an intermediary. In the first stage, an autoregressive transformer is used for the conditional generation of these style-rich tokens from text and control signals. The second stage generates codec tokens from both text and sampled style-rich tokens. Experiments show that training the first-stage model on extensive datasets enhances the content robustness of the two-stage model as well as control capabilities over multiple attributes. By selectively combining discrete labels and speaker embeddings, we explore fully controlling the speaker's timbre and other stylistic information, and adjusting attributes like emotion for a specified speaker. Audio samples are available at https://style-ar-tts.github.io.
title Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation
topic Multimedia
url https://arxiv.org/abs/2506.02997