DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Haoyang, Chen, Jun, Wu, Donghang, Li, Yuxin, Zhang, Yuxin, Zhang, Xiangyu Tony, Liu, Che, Lin, Qingjian, Peng, Yizhou, Liu, Hexin, Chng, Eng Siong, Yan, Chao, Wu, Boyong, Huang, Yechang, Yang, Xuerui, Tian, Fei
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911700612022272
author Zhang, Haoyang
Chen, Jun
Wu, Donghang
Li, Yuxin
Zhang, Yuxin
Zhang, Xiangyu Tony
Liu, Che
Lin, Qingjian
Peng, Yizhou
Liu, Hexin
Chng, Eng Siong
Yan, Chao
Wu, Boyong
Huang, Yechang
Yang, Xuerui
Tian, Fei
author_facet Zhang, Haoyang
Chen, Jun
Wu, Donghang
Li, Yuxin
Zhang, Yuxin
Zhang, Xiangyu Tony
Liu, Che
Lin, Qingjian
Peng, Yizhou
Liu, Hexin
Chng, Eng Siong
Yan, Chao
Wu, Boyong
Huang, Yechang
Yang, Xuerui
Tian, Fei
contents Recent advances in spoken dialogue language models have shifted from turn-based to full-duplex designs, where the model continuously listens to the user while generating responses. However, existing duplex backbones still lack a native channel for in-conversation planning and tool calling, leaving real-time agentic behaviour either tied to turn boundaries or relegated to an external cascade. We propose DuplexSLA, a native full-duplex Speech-Language-Action foundation model that decodes assistant audio together with a structured action stream on a shared 160 ms chunk timeline. DuplexSLA is built on a dual-stream three-channel formulation: a continuous user audio channel, a discrete assistant audio channel, and a rate-limited textual action channel, all decoded jointly by a single backbone, so that listening, speaking, planning, and tool calling unfold on one shared clock. Two capabilities define the model: (1) semantic-driven turn-taking control, where interruption, pause, and backchannel are handled inside the same backbone instead of by an external semantic VAD; and (2) in-conversation planning and tool calling, where planning text and structured tool calls are emitted on the action channel without halting assistant audio, so that multi-action and backchannel-triggered tool use are interleaved with ongoing speech. To evaluate these capabilities together, we further construct DuplexSLA-Bench, a duplex benchmark covering pause, interrupt, and backchannel turn-taking together with three styles of in-conversation tool calling. Our project page, interactive demos, and the DuplexSLA-Bench evaluation suite are publicly available at https://github.com/hyzhang24/DuplexSLA.
format Preprint
id arxiv_https___arxiv_org_abs_2605_20755
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action
Zhang, Haoyang
Chen, Jun
Wu, Donghang
Li, Yuxin
Zhang, Yuxin
Zhang, Xiangyu Tony
Liu, Che
Lin, Qingjian
Peng, Yizhou
Liu, Hexin
Chng, Eng Siong
Yan, Chao
Wu, Boyong
Huang, Yechang
Yang, Xuerui
Tian, Fei
Audio and Speech Processing
Recent advances in spoken dialogue language models have shifted from turn-based to full-duplex designs, where the model continuously listens to the user while generating responses. However, existing duplex backbones still lack a native channel for in-conversation planning and tool calling, leaving real-time agentic behaviour either tied to turn boundaries or relegated to an external cascade. We propose DuplexSLA, a native full-duplex Speech-Language-Action foundation model that decodes assistant audio together with a structured action stream on a shared 160 ms chunk timeline. DuplexSLA is built on a dual-stream three-channel formulation: a continuous user audio channel, a discrete assistant audio channel, and a rate-limited textual action channel, all decoded jointly by a single backbone, so that listening, speaking, planning, and tool calling unfold on one shared clock. Two capabilities define the model: (1) semantic-driven turn-taking control, where interruption, pause, and backchannel are handled inside the same backbone instead of by an external semantic VAD; and (2) in-conversation planning and tool calling, where planning text and structured tool calls are emitted on the action channel without halting assistant audio, so that multi-action and backchannel-triggered tool use are interleaved with ongoing speech. To evaluate these capabilities together, we further construct DuplexSLA-Bench, a duplex benchmark covering pause, interrupt, and backchannel turn-taking together with three styles of in-conversation tool calling. Our project page, interactive demos, and the DuplexSLA-Bench evaluation suite are publicly available at https://github.com/hyzhang24/DuplexSLA.
title DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action
topic Audio and Speech Processing
url https://arxiv.org/abs/2605.20755