What Makes a Good Speech Tokenizer for LLM-Centric Speech Generation? A Systematic Study

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Fan, Xiaoran, Sun, Zhichao, Gao, Yangfan, Xiong, Jingfei, Yan, Hang, Cao, Yifei, Sun, Jiajun, Li, Shuo, Zhang, Zhihao, Xi, Zhiheng, Zhou, Yuhao, Jin, Senjie, Jiang, Changhao, Ye, Junjie, Zhang, Ming, Zheng, Rui, Han, Zhenhua, Zhang, Yunke, Yan, Demei, Dong, Shaokang, Ji, Tao, Gui, Tao
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911379865206784
author Fan, Xiaoran
Sun, Zhichao
Gao, Yangfan
Xiong, Jingfei
Yan, Hang
Cao, Yifei
Sun, Jiajun
Li, Shuo
Zhang, Zhihao
Xi, Zhiheng
Zhou, Yuhao
Jin, Senjie
Jiang, Changhao
Ye, Junjie
Zhang, Ming
Zheng, Rui
Han, Zhenhua
Zhang, Yunke
Yan, Demei
Dong, Shaokang
Ji, Tao
Gui, Tao
author_facet Fan, Xiaoran
Sun, Zhichao
Gao, Yangfan
Xiong, Jingfei
Yan, Hang
Cao, Yifei
Sun, Jiajun
Li, Shuo
Zhang, Zhihao
Xi, Zhiheng
Zhou, Yuhao
Jin, Senjie
Jiang, Changhao
Ye, Junjie
Zhang, Ming
Zheng, Rui
Han, Zhenhua
Zhang, Yunke
Yan, Demei
Dong, Shaokang
Ji, Tao
Gui, Tao
contents Speech-language models (SLMs) offer a promising path toward unifying speech and text understanding and generation. However, challenges remain in achieving effective cross-modal alignment and high-quality speech generation. In this work, we systematically investigate the role of speech tokenizer designs in LLM-centric SLMs, augmented by speech heads and speaker modeling. We compare coupled, semi-decoupled, and fully decoupled speech tokenizers under a fair SLM framework and find that decoupled tokenization significantly improves alignment and synthesis quality. To address the information density mismatch between speech and text, we introduce multi-token prediction (MTP) into SLMs, enabling each hidden state to decode multiple speech tokens. This leads to up to 12$\times$ faster decoding and a substantial drop in word error rate (from 6.07 to 3.01). Furthermore, we propose a speaker-aware generation paradigm and introduce RoleTriviaQA, a large-scale role-playing knowledge QA benchmark with diverse speaker identities. Experiments demonstrate that our methods enhance both knowledge understanding and speaker consistency.
format Preprint
id arxiv_https___arxiv_org_abs_2506_12537
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle What Makes a Good Speech Tokenizer for LLM-Centric Speech Generation? A Systematic Study
Fan, Xiaoran
Sun, Zhichao
Gao, Yangfan
Xiong, Jingfei
Yan, Hang
Cao, Yifei
Sun, Jiajun
Li, Shuo
Zhang, Zhihao
Xi, Zhiheng
Zhou, Yuhao
Jin, Senjie
Jiang, Changhao
Ye, Junjie
Zhang, Ming
Zheng, Rui
Han, Zhenhua
Zhang, Yunke
Yan, Demei
Dong, Shaokang
Ji, Tao
Gui, Tao
Computation and Language
Artificial Intelligence
Audio and Speech Processing
Speech-language models (SLMs) offer a promising path toward unifying speech and text understanding and generation. However, challenges remain in achieving effective cross-modal alignment and high-quality speech generation. In this work, we systematically investigate the role of speech tokenizer designs in LLM-centric SLMs, augmented by speech heads and speaker modeling. We compare coupled, semi-decoupled, and fully decoupled speech tokenizers under a fair SLM framework and find that decoupled tokenization significantly improves alignment and synthesis quality. To address the information density mismatch between speech and text, we introduce multi-token prediction (MTP) into SLMs, enabling each hidden state to decode multiple speech tokens. This leads to up to 12$\times$ faster decoding and a substantial drop in word error rate (from 6.07 to 3.01). Furthermore, we propose a speaker-aware generation paradigm and introduce RoleTriviaQA, a large-scale role-playing knowledge QA benchmark with diverse speaker identities. Experiments demonstrate that our methods enhance both knowledge understanding and speaker consistency.
title What Makes a Good Speech Tokenizer for LLM-Centric Speech Generation? A Systematic Study
topic Computation and Language
Artificial Intelligence
Audio and Speech Processing
url https://arxiv.org/abs/2506.12537