SLM-SS: Speech Language Model for Generative Speech Separation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866910002161123328 |
|---|---|
| author | Li, Tianhua Li, Chenda Wang, Wei Zhou, Xin Chen, Xihui Gao, Jianqing Qian, Yanmin |
| author_facet | Li, Tianhua Li, Chenda Wang, Wei Zhou, Xin Chen, Xihui Gao, Jianqing Qian, Yanmin |
| contents | Speech separation (SS) has advanced significantly with neural network-based methods, showing improved performance on signal-level metrics. However, these methods often struggle to maintain speech intelligibility in the separated signals, which can negatively affect the performance of downstream tasks such as speech recognition. In this work, we propose SLM-SS, a novel approach that applies speech language models to SS, aiming to enhance the intelligibility and coherence of the separated signals. We frame SS as discrete multi-codebook sequence generation, using Encoder-Decoder models to map quantized speech mixtures to target tokens. In addition to the autoregressive modeling strategy, we introduce a non-autoregressive model to improve decoding efficiency for residual tokens. Experimental results on the LibriMix dataset demonstrate that our approach shows significantly better preservation of speech intelligibility, leading to improved linguistic consistency in a variety of downstream tasks compared to existing approaches. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2601_19533 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | SLM-SS: Speech Language Model for Generative Speech Separation Li, Tianhua Li, Chenda Wang, Wei Zhou, Xin Chen, Xihui Gao, Jianqing Qian, Yanmin Sound Artificial Intelligence Speech separation (SS) has advanced significantly with neural network-based methods, showing improved performance on signal-level metrics. However, these methods often struggle to maintain speech intelligibility in the separated signals, which can negatively affect the performance of downstream tasks such as speech recognition. In this work, we propose SLM-SS, a novel approach that applies speech language models to SS, aiming to enhance the intelligibility and coherence of the separated signals. We frame SS as discrete multi-codebook sequence generation, using Encoder-Decoder models to map quantized speech mixtures to target tokens. In addition to the autoregressive modeling strategy, we introduce a non-autoregressive model to improve decoding efficiency for residual tokens. Experimental results on the LibriMix dataset demonstrate that our approach shows significantly better preservation of speech intelligibility, leading to improved linguistic consistency in a variety of downstream tasks compared to existing approaches. |
| title | SLM-SS: Speech Language Model for Generative Speech Separation |
| topic | Sound Artificial Intelligence |
| url | https://arxiv.org/abs/2601.19533 |