SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Hu, Ke, Hosseini-Asl, Ehsan, Chen, Chen, Casanova, Edresson, Ghosh, Subhankar, Żelasko, Piotr, Chen, Zhehuai, Li, Jason, Balam, Jagadeesh, Ginsburg, Boris
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866911077052186624
author Hu, Ke
Hosseini-Asl, Ehsan
Chen, Chen
Casanova, Edresson
Ghosh, Subhankar
Żelasko, Piotr
Chen, Zhehuai
Li, Jason
Balam, Jagadeesh
Ginsburg, Boris
author_facet Hu, Ke
Hosseini-Asl, Ehsan
Chen, Chen
Casanova, Edresson
Ghosh, Subhankar
Żelasko, Piotr
Chen, Zhehuai
Li, Jason
Balam, Jagadeesh
Ginsburg, Boris
contents Spoken dialogue is an intuitive form of human-computer interaction, yet current speech language models often remain constrained to turn-based exchanges, lacking real-time adaptability such as user barge-in. We propose a novel duplex speech to speech (S2S) architecture featuring continuous user inputs and codec agent outputs with channel fusion that directly models simultaneous user and agent streams. Using a pretrained streaming encoder for user input enables the first duplex S2S model without requiring speech pretrain. Separate architectures for agent and user modeling facilitate codec fine-tuning for better agent voices and halve the bitrate (0.6 kbps) compared to previous works. Experimental results show that the proposed model outperforms previous duplex models in reasoning, turn-taking, and barge-in abilities. The model requires significantly less speech data, as speech pretrain is skipped, which markedly simplifies the process of building a duplex S2S model from any LLMs. Finally, it is the first openly available duplex S2S model with training and inference code to foster reproducibility.
format Preprint
id arxiv_https___arxiv_org_abs_2505_15670
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model
Hu, Ke
Hosseini-Asl, Ehsan
Chen, Chen
Casanova, Edresson
Ghosh, Subhankar
Żelasko, Piotr
Chen, Zhehuai
Li, Jason
Balam, Jagadeesh
Ginsburg, Boris
Computation and Language
Sound
Audio and Speech Processing
Spoken dialogue is an intuitive form of human-computer interaction, yet current speech language models often remain constrained to turn-based exchanges, lacking real-time adaptability such as user barge-in. We propose a novel duplex speech to speech (S2S) architecture featuring continuous user inputs and codec agent outputs with channel fusion that directly models simultaneous user and agent streams. Using a pretrained streaming encoder for user input enables the first duplex S2S model without requiring speech pretrain. Separate architectures for agent and user modeling facilitate codec fine-tuning for better agent voices and halve the bitrate (0.6 kbps) compared to previous works. Experimental results show that the proposed model outperforms previous duplex models in reasoning, turn-taking, and barge-in abilities. The model requires significantly less speech data, as speech pretrain is skipped, which markedly simplifies the process of building a duplex S2S model from any LLMs. Finally, it is the first openly available duplex S2S model with training and inference code to foster reproducibility.
title SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model
topic Computation and Language
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2505.15670