Comprehend and Talk: Text to Speech Synthesis via Dual Language Modeling

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cao, Junjie, Han, Yichen, Zhang, Ruonan, Hao, Xiaoyang, Li, Hongxiang, Zhao, Shuaijiang, Liu, Yue, Zhng, Xiao-Ping
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908560839933952
author Cao, Junjie
Han, Yichen
Zhang, Ruonan
Hao, Xiaoyang
Li, Hongxiang
Zhao, Shuaijiang
Liu, Yue
Zhng, Xiao-Ping
author_facet Cao, Junjie
Han, Yichen
Zhang, Ruonan
Hao, Xiaoyang
Li, Hongxiang
Zhao, Shuaijiang
Liu, Yue
Zhng, Xiao-Ping
contents Existing Large Language Model (LLM) based autoregressive (AR) text-to-speech (TTS) systems, while achieving state-of-the-art quality, still face critical challenges. The foundation of this LLM-based paradigm is the discretization of the continuous speech waveform into a sequence of discrete tokens by neural audio codec. However, single codebook modeling is well suited to text LLMs, but suffers from significant information loss; hierarchical acoustic tokens, typically generated via Residual Vector Quantization (RVQ), often lack explicit semantic structure, placing a heavy learning burden on the model. Furthermore, the autoregressive process is inherently susceptible to error accumulation, which can degrade generation stability. To address these limitations, we propose CaT-TTS, a novel framework for robust and semantically-grounded zero-shot synthesis. First, we introduce S3Codec, a split RVQ codec that injects explicit linguistic features into its primary codebook via semantic distillation from a state-of-the-art ASR model, providing a structured representation that simplifies the learning task. Second, we propose an ``Understand-then-Generate'' dual-Transformer architecture that decouples comprehension from rendering. An initial ``Understanding'' Transformer models the cross-modal relationship between text and the audio's semantic tokens to form a high-level utterance plan. A subsequent ``Generation'' Transformer then executes this plan, autoregressively synthesizing hierarchical acoustic tokens. Finally, to enhance generation stability, we introduce Masked Audio Parallel Inference (MAPI), a nearly parameter-free inference strategy that dynamically guides the decoding process to mitigate local errors.
format Preprint
id arxiv_https___arxiv_org_abs_2509_22062
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Comprehend and Talk: Text to Speech Synthesis via Dual Language Modeling
Cao, Junjie
Han, Yichen
Zhang, Ruonan
Hao, Xiaoyang
Li, Hongxiang
Zhao, Shuaijiang
Liu, Yue
Zhng, Xiao-Ping
Sound
Audio and Speech Processing
Existing Large Language Model (LLM) based autoregressive (AR) text-to-speech (TTS) systems, while achieving state-of-the-art quality, still face critical challenges. The foundation of this LLM-based paradigm is the discretization of the continuous speech waveform into a sequence of discrete tokens by neural audio codec. However, single codebook modeling is well suited to text LLMs, but suffers from significant information loss; hierarchical acoustic tokens, typically generated via Residual Vector Quantization (RVQ), often lack explicit semantic structure, placing a heavy learning burden on the model. Furthermore, the autoregressive process is inherently susceptible to error accumulation, which can degrade generation stability. To address these limitations, we propose CaT-TTS, a novel framework for robust and semantically-grounded zero-shot synthesis. First, we introduce S3Codec, a split RVQ codec that injects explicit linguistic features into its primary codebook via semantic distillation from a state-of-the-art ASR model, providing a structured representation that simplifies the learning task. Second, we propose an ``Understand-then-Generate'' dual-Transformer architecture that decouples comprehension from rendering. An initial ``Understanding'' Transformer models the cross-modal relationship between text and the audio's semantic tokens to form a high-level utterance plan. A subsequent ``Generation'' Transformer then executes this plan, autoregressively synthesizing hierarchical acoustic tokens. Finally, to enhance generation stability, we introduce Masked Audio Parallel Inference (MAPI), a nearly parameter-free inference strategy that dynamically guides the decoding process to mitigate local errors.
title Comprehend and Talk: Text to Speech Synthesis via Dual Language Modeling
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2509.22062