FireRedChat: A Pluggable, Full-Duplex Voice Interaction System with Cascaded and Semi-Cascaded Implementations

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Chen, Junjie, Hu, Yao, Li, Junjie, Li, Kangyue, Liu, Kun, Li, Wenpeng, Li, Xu, Li, Ziyuan, Shen, Feiyu, Tang, Xu, Wei, Manzhen, Wu, Yichen, Xie, Fenglong, Xu, Kaituo, Xie, Kun
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912576235896832
author Chen, Junjie
Hu, Yao
Li, Junjie
Li, Kangyue
Liu, Kun
Li, Wenpeng
Li, Xu
Li, Ziyuan
Shen, Feiyu
Tang, Xu
Wei, Manzhen
Wu, Yichen
Xie, Fenglong
Xu, Kaituo
Xie, Kun
author_facet Chen, Junjie
Hu, Yao
Li, Junjie
Li, Kangyue
Liu, Kun
Li, Wenpeng
Li, Xu
Li, Ziyuan
Shen, Feiyu
Tang, Xu
Wei, Manzhen
Wu, Yichen
Xie, Fenglong
Xu, Kaituo
Xie, Kun
contents Full-duplex voice interaction allows users and agents to speak simultaneously with controllable barge-in, enabling lifelike assistants and customer service. Existing solutions are either end-to-end, difficult to design and hard to control, or modular pipelines governed by turn-taking controllers that ease upgrades and per-module optimization; however, prior modular frameworks depend on non-open components and external providers, limiting holistic optimization. In this work, we present a complete, practical full-duplex voice interaction system comprising a turn-taking controller, an interaction module, and a dialogue manager. The controller integrates streaming personalized VAD (pVAD) to suppress false barge-ins from noise and non-primary speakers, precisely timestamp primary-speaker segments, and explicitly enable primary-speaker barge-ins; a semantic end-of-turn detector improves stop decisions. It upgrades heterogeneous half-duplex pipelines, cascaded, semi-cascaded, and speech-to-speech, to full duplex. Using internal models, we implement cascaded and semi-cascaded variants; the semi-cascaded one captures emotional and paralinguistic cues, yields more coherent responses, lowers latency and error propagation, and improves robustness. A dialogue manager extends capabilities via tool invocation and context management. We also propose three system-level metrics, barge-in, end-of-turn detection accuracy, and end-to-end latency, to assess naturalness, control accuracy, and efficiency. Experiments show fewer false interruptions, more accurate semantic ends, and lower latency approaching industrial systems, enabling robust, natural, real-time full-duplex interaction. Demos: https://fireredteam.github.io/demos/firered_chat.
format Preprint
id arxiv_https___arxiv_org_abs_2509_06502
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle FireRedChat: A Pluggable, Full-Duplex Voice Interaction System with Cascaded and Semi-Cascaded Implementations
Chen, Junjie
Hu, Yao
Li, Junjie
Li, Kangyue
Liu, Kun
Li, Wenpeng
Li, Xu
Li, Ziyuan
Shen, Feiyu
Tang, Xu
Wei, Manzhen
Wu, Yichen
Xie, Fenglong
Xu, Kaituo
Xie, Kun
Sound
Human-Computer Interaction
Full-duplex voice interaction allows users and agents to speak simultaneously with controllable barge-in, enabling lifelike assistants and customer service. Existing solutions are either end-to-end, difficult to design and hard to control, or modular pipelines governed by turn-taking controllers that ease upgrades and per-module optimization; however, prior modular frameworks depend on non-open components and external providers, limiting holistic optimization. In this work, we present a complete, practical full-duplex voice interaction system comprising a turn-taking controller, an interaction module, and a dialogue manager. The controller integrates streaming personalized VAD (pVAD) to suppress false barge-ins from noise and non-primary speakers, precisely timestamp primary-speaker segments, and explicitly enable primary-speaker barge-ins; a semantic end-of-turn detector improves stop decisions. It upgrades heterogeneous half-duplex pipelines, cascaded, semi-cascaded, and speech-to-speech, to full duplex. Using internal models, we implement cascaded and semi-cascaded variants; the semi-cascaded one captures emotional and paralinguistic cues, yields more coherent responses, lowers latency and error propagation, and improves robustness. A dialogue manager extends capabilities via tool invocation and context management. We also propose three system-level metrics, barge-in, end-of-turn detection accuracy, and end-to-end latency, to assess naturalness, control accuracy, and efficiency. Experiments show fewer false interruptions, more accurate semantic ends, and lower latency approaching industrial systems, enabling robust, natural, real-time full-duplex interaction. Demos: https://fireredteam.github.io/demos/firered_chat.
title FireRedChat: A Pluggable, Full-Duplex Voice Interaction System with Cascaded and Semi-Cascaded Implementations
topic Sound
Human-Computer Interaction
url https://arxiv.org/abs/2509.06502