Balancing Speech Understanding and Generation Using Continual Pre-training for Codec-based Speech LLM

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shi, Jiatong, Zhang, Chunlei, Tian, Jinchuan, Ni, Junrui, Zhang, Hao, Watanabe, Shinji, Yu, Dong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909929504243712
author Shi, Jiatong
Zhang, Chunlei
Tian, Jinchuan
Ni, Junrui
Zhang, Hao
Watanabe, Shinji
Yu, Dong
author_facet Shi, Jiatong
Zhang, Chunlei
Tian, Jinchuan
Ni, Junrui
Zhang, Hao
Watanabe, Shinji
Yu, Dong
contents Recent advances in speech language models (LLMs) have extended textual LLMs to the speech domain, but balancing speech understanding and generation remains challenging, especially with codec-based representations. We propose a continual pre-training (CPT) framework that adapts a textual LLM to handle codec-discretized speech, mitigating modality mismatch and preserving linguistic reasoning. Our unified model supports both understanding and generation, achieving strong results across ASR, TTS, S2T-Trans, and S2S-Trans. Notably, we present the first end-to-end, single-pass S2S-Trans system using only neural codec tokens, without intermediate transcriptions, translations, or semantic tokens. CPT proves essential for cross-modal alignment and task generalization, making it a powerful tool for building robust, unified speech LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2502_16897
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Balancing Speech Understanding and Generation Using Continual Pre-training for Codec-based Speech LLM
Shi, Jiatong
Zhang, Chunlei
Tian, Jinchuan
Ni, Junrui
Zhang, Hao
Watanabe, Shinji
Yu, Dong
Audio and Speech Processing
Recent advances in speech language models (LLMs) have extended textual LLMs to the speech domain, but balancing speech understanding and generation remains challenging, especially with codec-based representations. We propose a continual pre-training (CPT) framework that adapts a textual LLM to handle codec-discretized speech, mitigating modality mismatch and preserving linguistic reasoning. Our unified model supports both understanding and generation, achieving strong results across ASR, TTS, S2T-Trans, and S2S-Trans. Notably, we present the first end-to-end, single-pass S2S-Trans system using only neural codec tokens, without intermediate transcriptions, translations, or semantic tokens. CPT proves essential for cross-modal alignment and task generalization, making it a powerful tool for building robust, unified speech LLMs.
title Balancing Speech Understanding and Generation Using Continual Pre-training for Codec-based Speech LLM
topic Audio and Speech Processing
url https://arxiv.org/abs/2502.16897