SLM-SS: Speech Language Model for Generative Speech Separation

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Li, Tianhua, Li, Chenda, Wang, Wei, Zhou, Xin, Chen, Xihui, Gao, Jianqing, Qian, Yanmin
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866910002161123328
author Li, Tianhua
Li, Chenda
Wang, Wei
Zhou, Xin
Chen, Xihui
Gao, Jianqing
Qian, Yanmin
author_facet Li, Tianhua
Li, Chenda
Wang, Wei
Zhou, Xin
Chen, Xihui
Gao, Jianqing
Qian, Yanmin
contents Speech separation (SS) has advanced significantly with neural network-based methods, showing improved performance on signal-level metrics. However, these methods often struggle to maintain speech intelligibility in the separated signals, which can negatively affect the performance of downstream tasks such as speech recognition. In this work, we propose SLM-SS, a novel approach that applies speech language models to SS, aiming to enhance the intelligibility and coherence of the separated signals. We frame SS as discrete multi-codebook sequence generation, using Encoder-Decoder models to map quantized speech mixtures to target tokens. In addition to the autoregressive modeling strategy, we introduce a non-autoregressive model to improve decoding efficiency for residual tokens. Experimental results on the LibriMix dataset demonstrate that our approach shows significantly better preservation of speech intelligibility, leading to improved linguistic consistency in a variety of downstream tasks compared to existing approaches.
format Preprint
id arxiv_https___arxiv_org_abs_2601_19533
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle SLM-SS: Speech Language Model for Generative Speech Separation
Li, Tianhua
Li, Chenda
Wang, Wei
Zhou, Xin
Chen, Xihui
Gao, Jianqing
Qian, Yanmin
Sound
Artificial Intelligence
Speech separation (SS) has advanced significantly with neural network-based methods, showing improved performance on signal-level metrics. However, these methods often struggle to maintain speech intelligibility in the separated signals, which can negatively affect the performance of downstream tasks such as speech recognition. In this work, we propose SLM-SS, a novel approach that applies speech language models to SS, aiming to enhance the intelligibility and coherence of the separated signals. We frame SS as discrete multi-codebook sequence generation, using Encoder-Decoder models to map quantized speech mixtures to target tokens. In addition to the autoregressive modeling strategy, we introduce a non-autoregressive model to improve decoding efficiency for residual tokens. Experimental results on the LibriMix dataset demonstrate that our approach shows significantly better preservation of speech intelligibility, leading to improved linguistic consistency in a variety of downstream tasks compared to existing approaches.
title SLM-SS: Speech Language Model for Generative Speech Separation
topic Sound
Artificial Intelligence
url https://arxiv.org/abs/2601.19533