GOAT-TTS: Expressive and Realistic Speech Generation via A Dual-Branch LLM

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Song, Yaodong, Chen, Hongjie, Lian, Jie, Zhang, Yuxin, Xia, Guangmin, Li, Zehan, Zhao, Genliang, Kang, Jian, Li, Jie, Li, Yongxiang, Li, Xuelong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908382553702400
author Song, Yaodong
Chen, Hongjie
Lian, Jie
Zhang, Yuxin
Xia, Guangmin
Li, Zehan
Zhao, Genliang
Kang, Jian
Li, Jie
Li, Yongxiang
Li, Xuelong
author_facet Song, Yaodong
Chen, Hongjie
Lian, Jie
Zhang, Yuxin
Xia, Guangmin
Li, Zehan
Zhao, Genliang
Kang, Jian
Li, Jie
Li, Yongxiang
Li, Xuelong
contents While large language models (LLMs) have revolutionized text-to-speech (TTS) synthesis through discrete tokenization paradigms, current architectures exhibit fundamental tensions between three critical dimensions: 1) irreversible loss of acoustic characteristics caused by quantization of speech prompts; 2) stringent dependence on precisely aligned prompt speech-text pairs that limit real-world deployment; and 3) catastrophic forgetting of the LLM's native text comprehension during optimization for speech token generation. To address these challenges, we propose an LLM-based text-to-speech Generation approach Optimized via a novel dual-branch ArchiTecture (GOAT-TTS). Our framework introduces two key innovations: (1) The modality-alignment branch combines a speech encoder and projector to capture continuous acoustic embeddings, enabling bidirectional correlation between paralinguistic features (language, timbre, emotion) and semantic text representations without transcript dependency; (2) The speech-generation branch employs modular fine-tuning on top-k layers of an LLM for speech token prediction while freezing the bottom-n layers to preserve foundational linguistic knowledge. Moreover, multi-token prediction is introduced to support real-time streaming TTS synthesis. Experimental results demonstrate that our GOAT-TTS achieves performance comparable to state-of-the-art TTS models while validating the efficacy of synthesized dialect speech data.
format Preprint
id arxiv_https___arxiv_org_abs_2504_12339
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle GOAT-TTS: Expressive and Realistic Speech Generation via A Dual-Branch LLM
Song, Yaodong
Chen, Hongjie
Lian, Jie
Zhang, Yuxin
Xia, Guangmin
Li, Zehan
Zhao, Genliang
Kang, Jian
Li, Jie
Li, Yongxiang
Li, Xuelong
Computation and Language
Sound
Audio and Speech Processing
While large language models (LLMs) have revolutionized text-to-speech (TTS) synthesis through discrete tokenization paradigms, current architectures exhibit fundamental tensions between three critical dimensions: 1) irreversible loss of acoustic characteristics caused by quantization of speech prompts; 2) stringent dependence on precisely aligned prompt speech-text pairs that limit real-world deployment; and 3) catastrophic forgetting of the LLM's native text comprehension during optimization for speech token generation. To address these challenges, we propose an LLM-based text-to-speech Generation approach Optimized via a novel dual-branch ArchiTecture (GOAT-TTS). Our framework introduces two key innovations: (1) The modality-alignment branch combines a speech encoder and projector to capture continuous acoustic embeddings, enabling bidirectional correlation between paralinguistic features (language, timbre, emotion) and semantic text representations without transcript dependency; (2) The speech-generation branch employs modular fine-tuning on top-k layers of an LLM for speech token prediction while freezing the bottom-n layers to preserve foundational linguistic knowledge. Moreover, multi-token prediction is introduced to support real-time streaming TTS synthesis. Experimental results demonstrate that our GOAT-TTS achieves performance comparable to state-of-the-art TTS models while validating the efficacy of synthesized dialect speech data.
title GOAT-TTS: Expressive and Realistic Speech Generation via A Dual-Branch LLM
topic Computation and Language
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2504.12339