Balancing Speech Understanding and Generation Using Continual Pre-training for Codec-based Speech LLM
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866909929504243712 |
|---|---|
| author | Shi, Jiatong Zhang, Chunlei Tian, Jinchuan Ni, Junrui Zhang, Hao Watanabe, Shinji Yu, Dong |
| author_facet | Shi, Jiatong Zhang, Chunlei Tian, Jinchuan Ni, Junrui Zhang, Hao Watanabe, Shinji Yu, Dong |
| contents | Recent advances in speech language models (LLMs) have extended textual LLMs to the speech domain, but balancing speech understanding and generation remains challenging, especially with codec-based representations. We propose a continual pre-training (CPT) framework that adapts a textual LLM to handle codec-discretized speech, mitigating modality mismatch and preserving linguistic reasoning. Our unified model supports both understanding and generation, achieving strong results across ASR, TTS, S2T-Trans, and S2S-Trans. Notably, we present the first end-to-end, single-pass S2S-Trans system using only neural codec tokens, without intermediate transcriptions, translations, or semantic tokens. CPT proves essential for cross-modal alignment and task generalization, making it a powerful tool for building robust, unified speech LLMs. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2502_16897 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Balancing Speech Understanding and Generation Using Continual Pre-training for Codec-based Speech LLM Shi, Jiatong Zhang, Chunlei Tian, Jinchuan Ni, Junrui Zhang, Hao Watanabe, Shinji Yu, Dong Audio and Speech Processing Recent advances in speech language models (LLMs) have extended textual LLMs to the speech domain, but balancing speech understanding and generation remains challenging, especially with codec-based representations. We propose a continual pre-training (CPT) framework that adapts a textual LLM to handle codec-discretized speech, mitigating modality mismatch and preserving linguistic reasoning. Our unified model supports both understanding and generation, achieving strong results across ASR, TTS, S2T-Trans, and S2S-Trans. Notably, we present the first end-to-end, single-pass S2S-Trans system using only neural codec tokens, without intermediate transcriptions, translations, or semantic tokens. CPT proves essential for cross-modal alignment and task generalization, making it a powerful tool for building robust, unified speech LLMs. |
| title | Balancing Speech Understanding and Generation Using Continual Pre-training for Codec-based Speech LLM |
| topic | Audio and Speech Processing |
| url | https://arxiv.org/abs/2502.16897 |