Improving endpoint detection in end-to-end streaming ASR for conversational speech

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: C, Anandh, Durai, Karthik Pandia, Prakash, Jeena, Arumugam, Manickavela, Hacioglu, Kadri, Dubagunta, S. Pavankumar, Stolcke, Andreas, Venkatesan, Shankar, Ganapathiraju, Aravind
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913854639833088
author C, Anandh
Durai, Karthik Pandia
Prakash, Jeena
Arumugam, Manickavela
Hacioglu, Kadri
Dubagunta, S. Pavankumar
Stolcke, Andreas
Venkatesan, Shankar
Ganapathiraju, Aravind
author_facet C, Anandh
Durai, Karthik Pandia
Prakash, Jeena
Arumugam, Manickavela
Hacioglu, Kadri
Dubagunta, S. Pavankumar
Stolcke, Andreas
Venkatesan, Shankar
Ganapathiraju, Aravind
contents ASR endpointing (EP) plays a major role in delivering a good user experience in products supporting human or artificial agents in human-human/machine conversations. Transducer-based ASR (T-ASR) is an end-to-end (E2E) ASR modelling technique preferred for streaming. A major limitation of T-ASR is delayed emission of ASR outputs, which could lead to errors or delays in EP. Inaccurate EP will cut the user off while speaking, returning incomplete transcript while delays in EP will increase the perceived latency, degrading the user experience. We propose methods to improve EP by addressing delayed emission along with EP mistakes. To address the delayed emission problem, we introduce an end-of-word token at the end of each word, along with a delay penalty. The EP delay is addressed by obtaining a reliable frame-level speech activity detection using an auxiliary network. We apply the proposed methods on Switchboard conversational speech corpus and evaluate it against a delay penalty method.
format Preprint
id arxiv_https___arxiv_org_abs_2505_17070
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Improving endpoint detection in end-to-end streaming ASR for conversational speech
C, Anandh
Durai, Karthik Pandia
Prakash, Jeena
Arumugam, Manickavela
Hacioglu, Kadri
Dubagunta, S. Pavankumar
Stolcke, Andreas
Venkatesan, Shankar
Ganapathiraju, Aravind
Computation and Language
Artificial Intelligence
Sound
Audio and Speech Processing
ASR endpointing (EP) plays a major role in delivering a good user experience in products supporting human or artificial agents in human-human/machine conversations. Transducer-based ASR (T-ASR) is an end-to-end (E2E) ASR modelling technique preferred for streaming. A major limitation of T-ASR is delayed emission of ASR outputs, which could lead to errors or delays in EP. Inaccurate EP will cut the user off while speaking, returning incomplete transcript while delays in EP will increase the perceived latency, degrading the user experience. We propose methods to improve EP by addressing delayed emission along with EP mistakes. To address the delayed emission problem, we introduce an end-of-word token at the end of each word, along with a delay penalty. The EP delay is addressed by obtaining a reliable frame-level speech activity detection using an auxiliary network. We apply the proposed methods on Switchboard conversational speech corpus and evaluate it against a delay penalty method.
title Improving endpoint detection in end-to-end streaming ASR for conversational speech
topic Computation and Language
Artificial Intelligence
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2505.17070