Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Udupa, Sathvik, Watanabe, Shinji, Schwarz, Petr, Cernocky, Jan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909652836417536
author Udupa, Sathvik
Watanabe, Shinji
Schwarz, Petr
Cernocky, Jan
author_facet Udupa, Sathvik
Watanabe, Shinji
Schwarz, Petr
Cernocky, Jan
contents Accurate, low-latency endpointing is crucial for effective spoken dialogue systems. While traditional endpointers often rely on spectrum-based audio features, this work proposes real-time speech endpointing for multi-turn dialogues using streaming, low-bitrate Neural Audio Codec (NAC) features, building upon recent advancements in neural audio codecs. To further reduce cutoff errors, we introduce a novel label delay training scheme. At a fixed median latency of 160 ms, our combined NAC and label delay approach achieves significant relative cutoff error reductions: 42.7% for a single-stream endpointer and 37.5% for a two-stream configuration, compared to baseline methods. Finally, we demonstrate efficient integration with a codec-based pretrained speech large language model, improving its median response time by 1200 ms and reducing its cutoff error by 35%.
format Preprint
id arxiv_https___arxiv_org_abs_2506_07081
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training
Udupa, Sathvik
Watanabe, Shinji
Schwarz, Petr
Cernocky, Jan
Sound
Audio and Speech Processing
Accurate, low-latency endpointing is crucial for effective spoken dialogue systems. While traditional endpointers often rely on spectrum-based audio features, this work proposes real-time speech endpointing for multi-turn dialogues using streaming, low-bitrate Neural Audio Codec (NAC) features, building upon recent advancements in neural audio codecs. To further reduce cutoff errors, we introduce a novel label delay training scheme. At a fixed median latency of 160 ms, our combined NAC and label delay approach achieves significant relative cutoff error reductions: 42.7% for a single-stream endpointer and 37.5% for a two-stream configuration, compared to baseline methods. Finally, we demonstrate efficient integration with a codec-based pretrained speech large language model, improving its median response time by 1200 ms and reducing its cutoff error by 35%.
title Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2506.07081