Phoenix-VAD: Streaming Semantic Endpoint Detection for Full-Duplex Speech Interaction
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866912686360494080 |
|---|---|
| author | Wu, Weijie Guan, Wenhao Wang, Kaidi Chen, Peijie Zha, Zhuanling Li, Junbo Fang, Jun Li, Lin Hong, Qingyang |
| author_facet | Wu, Weijie Guan, Wenhao Wang, Kaidi Chen, Peijie Zha, Zhuanling Li, Junbo Fang, Jun Li, Lin Hong, Qingyang |
| contents | Spoken dialogue models have significantly advanced intelligent human-computer interaction, yet they lack a plug-and-play full-duplex prediction module for semantic endpoint detection, hindering seamless audio interactions. In this paper, we introduce Phoenix-VAD, an LLM-based model that enables streaming semantic endpoint detection. Specifically, Phoenix-VAD leverages the semantic comprehension capability of the LLM and a sliding window training strategy to achieve reliable semantic endpoint detection while supporting streaming inference. Experiments on both semantically complete and incomplete speech scenarios indicate that Phoenix-VAD achieves excellent and competitive performance. Furthermore, this design enables the full-duplex prediction module to be optimized independently of the dialogue model, providing more reliable and flexible support for next-generation human-computer interaction. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2509_20410 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Phoenix-VAD: Streaming Semantic Endpoint Detection for Full-Duplex Speech Interaction Wu, Weijie Guan, Wenhao Wang, Kaidi Chen, Peijie Zha, Zhuanling Li, Junbo Fang, Jun Li, Lin Hong, Qingyang Audio and Speech Processing Sound Spoken dialogue models have significantly advanced intelligent human-computer interaction, yet they lack a plug-and-play full-duplex prediction module for semantic endpoint detection, hindering seamless audio interactions. In this paper, we introduce Phoenix-VAD, an LLM-based model that enables streaming semantic endpoint detection. Specifically, Phoenix-VAD leverages the semantic comprehension capability of the LLM and a sliding window training strategy to achieve reliable semantic endpoint detection while supporting streaming inference. Experiments on both semantically complete and incomplete speech scenarios indicate that Phoenix-VAD achieves excellent and competitive performance. Furthermore, this design enables the full-duplex prediction module to be optimized independently of the dialogue model, providing more reliable and flexible support for next-generation human-computer interaction. |
| title | Phoenix-VAD: Streaming Semantic Endpoint Detection for Full-Duplex Speech Interaction |
| topic | Audio and Speech Processing Sound |
| url | https://arxiv.org/abs/2509.20410 |