Phoenix-VAD: Streaming Semantic Endpoint Detection for Full-Duplex Speech Interaction

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wu, Weijie, Guan, Wenhao, Wang, Kaidi, Chen, Peijie, Zha, Zhuanling, Li, Junbo, Fang, Jun, Li, Lin, Hong, Qingyang
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912686360494080
author Wu, Weijie
Guan, Wenhao
Wang, Kaidi
Chen, Peijie
Zha, Zhuanling
Li, Junbo
Fang, Jun
Li, Lin
Hong, Qingyang
author_facet Wu, Weijie
Guan, Wenhao
Wang, Kaidi
Chen, Peijie
Zha, Zhuanling
Li, Junbo
Fang, Jun
Li, Lin
Hong, Qingyang
contents Spoken dialogue models have significantly advanced intelligent human-computer interaction, yet they lack a plug-and-play full-duplex prediction module for semantic endpoint detection, hindering seamless audio interactions. In this paper, we introduce Phoenix-VAD, an LLM-based model that enables streaming semantic endpoint detection. Specifically, Phoenix-VAD leverages the semantic comprehension capability of the LLM and a sliding window training strategy to achieve reliable semantic endpoint detection while supporting streaming inference. Experiments on both semantically complete and incomplete speech scenarios indicate that Phoenix-VAD achieves excellent and competitive performance. Furthermore, this design enables the full-duplex prediction module to be optimized independently of the dialogue model, providing more reliable and flexible support for next-generation human-computer interaction.
format Preprint
id arxiv_https___arxiv_org_abs_2509_20410
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Phoenix-VAD: Streaming Semantic Endpoint Detection for Full-Duplex Speech Interaction
Wu, Weijie
Guan, Wenhao
Wang, Kaidi
Chen, Peijie
Zha, Zhuanling
Li, Junbo
Fang, Jun
Li, Lin
Hong, Qingyang
Audio and Speech Processing
Sound
Spoken dialogue models have significantly advanced intelligent human-computer interaction, yet they lack a plug-and-play full-duplex prediction module for semantic endpoint detection, hindering seamless audio interactions. In this paper, we introduce Phoenix-VAD, an LLM-based model that enables streaming semantic endpoint detection. Specifically, Phoenix-VAD leverages the semantic comprehension capability of the LLM and a sliding window training strategy to achieve reliable semantic endpoint detection while supporting streaming inference. Experiments on both semantically complete and incomplete speech scenarios indicate that Phoenix-VAD achieves excellent and competitive performance. Furthermore, this design enables the full-duplex prediction module to be optimized independently of the dialogue model, providing more reliable and flexible support for next-generation human-computer interaction.
title Phoenix-VAD: Streaming Semantic Endpoint Detection for Full-Duplex Speech Interaction
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2509.20410