Easy Turn: Integrating Acoustic and Linguistic Modalities for Robust Turn-Taking in Full-Duplex Spoken Dialogue Systems

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Guojian, Wang, Chengyou, Xue, Hongfei, Wang, Shuiyuan, Gao, Dehui, Zhang, Zihan, Lin, Yuke, Li, Wenjie, Xiao, Longshuai, Fu, Zhonghua, Xie, Lei
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918150001393664
author Li, Guojian
Wang, Chengyou
Xue, Hongfei
Wang, Shuiyuan
Gao, Dehui
Zhang, Zihan
Lin, Yuke
Li, Wenjie
Xiao, Longshuai
Fu, Zhonghua
Xie, Lei
author_facet Li, Guojian
Wang, Chengyou
Xue, Hongfei
Wang, Shuiyuan
Gao, Dehui
Zhang, Zihan
Lin, Yuke
Li, Wenjie
Xiao, Longshuai
Fu, Zhonghua
Xie, Lei
contents Full-duplex interaction is crucial for natural human-machine communication, yet remains challenging as it requires robust turn-taking detection to decide when the system should speak, listen, or remain silent. Existing solutions either rely on dedicated turn-taking models, most of which are not open-sourced. The few available ones are limited by their large parameter size or by supporting only a single modality, such as acoustic or linguistic. Alternatively, some approaches finetune LLM backbones to enable full-duplex capability, but this requires large amounts of full-duplex data, which remain scarce in open-source form. To address these issues, we propose Easy Turn, an open-source, modular turn-taking detection model that integrates acoustic and linguistic bimodal information to predict four dialogue turn states: complete, incomplete, backchannel, and wait, accompanied by the release of Easy Turn trainset, a 1,145-hour speech dataset designed for training turn-taking detection models. Compared to existing open-source models like TEN Turn Detection and Smart Turn V2, our model achieves state-of-the-art turn-taking detection accuracy on our open-source Easy Turn testset. The data and model will be made publicly available on GitHub.
format Preprint
id arxiv_https___arxiv_org_abs_2509_23938
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Easy Turn: Integrating Acoustic and Linguistic Modalities for Robust Turn-Taking in Full-Duplex Spoken Dialogue Systems
Li, Guojian
Wang, Chengyou
Xue, Hongfei
Wang, Shuiyuan
Gao, Dehui
Zhang, Zihan
Lin, Yuke
Li, Wenjie
Xiao, Longshuai
Fu, Zhonghua
Xie, Lei
Computation and Language
Artificial Intelligence
Full-duplex interaction is crucial for natural human-machine communication, yet remains challenging as it requires robust turn-taking detection to decide when the system should speak, listen, or remain silent. Existing solutions either rely on dedicated turn-taking models, most of which are not open-sourced. The few available ones are limited by their large parameter size or by supporting only a single modality, such as acoustic or linguistic. Alternatively, some approaches finetune LLM backbones to enable full-duplex capability, but this requires large amounts of full-duplex data, which remain scarce in open-source form. To address these issues, we propose Easy Turn, an open-source, modular turn-taking detection model that integrates acoustic and linguistic bimodal information to predict four dialogue turn states: complete, incomplete, backchannel, and wait, accompanied by the release of Easy Turn trainset, a 1,145-hour speech dataset designed for training turn-taking detection models. Compared to existing open-source models like TEN Turn Detection and Smart Turn V2, our model achieves state-of-the-art turn-taking detection accuracy on our open-source Easy Turn testset. The data and model will be made publicly available on GitHub.
title Easy Turn: Integrating Acoustic and Linguistic Modalities for Robust Turn-Taking in Full-Duplex Spoken Dialogue Systems
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2509.23938