What Makes a Good Speech Tokenizer for LLM-Centric Speech Generation? A Systematic Study
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866911379865206784 |
|---|---|
| author | Fan, Xiaoran Sun, Zhichao Gao, Yangfan Xiong, Jingfei Yan, Hang Cao, Yifei Sun, Jiajun Li, Shuo Zhang, Zhihao Xi, Zhiheng Zhou, Yuhao Jin, Senjie Jiang, Changhao Ye, Junjie Zhang, Ming Zheng, Rui Han, Zhenhua Zhang, Yunke Yan, Demei Dong, Shaokang Ji, Tao Gui, Tao |
| author_facet | Fan, Xiaoran Sun, Zhichao Gao, Yangfan Xiong, Jingfei Yan, Hang Cao, Yifei Sun, Jiajun Li, Shuo Zhang, Zhihao Xi, Zhiheng Zhou, Yuhao Jin, Senjie Jiang, Changhao Ye, Junjie Zhang, Ming Zheng, Rui Han, Zhenhua Zhang, Yunke Yan, Demei Dong, Shaokang Ji, Tao Gui, Tao |
| contents | Speech-language models (SLMs) offer a promising path toward unifying speech and text understanding and generation. However, challenges remain in achieving effective cross-modal alignment and high-quality speech generation. In this work, we systematically investigate the role of speech tokenizer designs in LLM-centric SLMs, augmented by speech heads and speaker modeling. We compare coupled, semi-decoupled, and fully decoupled speech tokenizers under a fair SLM framework and find that decoupled tokenization significantly improves alignment and synthesis quality. To address the information density mismatch between speech and text, we introduce multi-token prediction (MTP) into SLMs, enabling each hidden state to decode multiple speech tokens. This leads to up to 12$\times$ faster decoding and a substantial drop in word error rate (from 6.07 to 3.01). Furthermore, we propose a speaker-aware generation paradigm and introduce RoleTriviaQA, a large-scale role-playing knowledge QA benchmark with diverse speaker identities. Experiments demonstrate that our methods enhance both knowledge understanding and speaker consistency. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2506_12537 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | What Makes a Good Speech Tokenizer for LLM-Centric Speech Generation? A Systematic Study Fan, Xiaoran Sun, Zhichao Gao, Yangfan Xiong, Jingfei Yan, Hang Cao, Yifei Sun, Jiajun Li, Shuo Zhang, Zhihao Xi, Zhiheng Zhou, Yuhao Jin, Senjie Jiang, Changhao Ye, Junjie Zhang, Ming Zheng, Rui Han, Zhenhua Zhang, Yunke Yan, Demei Dong, Shaokang Ji, Tao Gui, Tao Computation and Language Artificial Intelligence Audio and Speech Processing Speech-language models (SLMs) offer a promising path toward unifying speech and text understanding and generation. However, challenges remain in achieving effective cross-modal alignment and high-quality speech generation. In this work, we systematically investigate the role of speech tokenizer designs in LLM-centric SLMs, augmented by speech heads and speaker modeling. We compare coupled, semi-decoupled, and fully decoupled speech tokenizers under a fair SLM framework and find that decoupled tokenization significantly improves alignment and synthesis quality. To address the information density mismatch between speech and text, we introduce multi-token prediction (MTP) into SLMs, enabling each hidden state to decode multiple speech tokens. This leads to up to 12$\times$ faster decoding and a substantial drop in word error rate (from 6.07 to 3.01). Furthermore, we propose a speaker-aware generation paradigm and introduce RoleTriviaQA, a large-scale role-playing knowledge QA benchmark with diverse speaker identities. Experiments demonstrate that our methods enhance both knowledge understanding and speaker consistency. |
| title | What Makes a Good Speech Tokenizer for LLM-Centric Speech Generation? A Systematic Study |
| topic | Computation and Language Artificial Intelligence Audio and Speech Processing |
| url | https://arxiv.org/abs/2506.12537 |