Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training
Fuente:
arXiv
Saved in:
| Main Authors: | , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866909652836417536 |
|---|---|
| author | Udupa, Sathvik Watanabe, Shinji Schwarz, Petr Cernocky, Jan |
| author_facet | Udupa, Sathvik Watanabe, Shinji Schwarz, Petr Cernocky, Jan |
| contents | Accurate, low-latency endpointing is crucial for effective spoken dialogue systems. While traditional endpointers often rely on spectrum-based audio features, this work proposes real-time speech endpointing for multi-turn dialogues using streaming, low-bitrate Neural Audio Codec (NAC) features, building upon recent advancements in neural audio codecs. To further reduce cutoff errors, we introduce a novel label delay training scheme. At a fixed median latency of 160 ms, our combined NAC and label delay approach achieves significant relative cutoff error reductions: 42.7% for a single-stream endpointer and 37.5% for a two-stream configuration, compared to baseline methods. Finally, we demonstrate efficient integration with a codec-based pretrained speech large language model, improving its median response time by 1200 ms and reducing its cutoff error by 35%. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2506_07081 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Udupa, Sathvik Watanabe, Shinji Schwarz, Petr Cernocky, Jan Sound Audio and Speech Processing Accurate, low-latency endpointing is crucial for effective spoken dialogue systems. While traditional endpointers often rely on spectrum-based audio features, this work proposes real-time speech endpointing for multi-turn dialogues using streaming, low-bitrate Neural Audio Codec (NAC) features, building upon recent advancements in neural audio codecs. To further reduce cutoff errors, we introduce a novel label delay training scheme. At a fixed median latency of 160 ms, our combined NAC and label delay approach achieves significant relative cutoff error reductions: 42.7% for a single-stream endpointer and 37.5% for a two-stream configuration, compared to baseline methods. Finally, we demonstrate efficient integration with a codec-based pretrained speech large language model, improving its median response time by 1200 ms and reducing its cutoff error by 35%. |
| title | Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training |
| topic | Sound Audio and Speech Processing |
| url | https://arxiv.org/abs/2506.07081 |