Contextual Biasing for Streaming ASR via CTC-based Word Spotting

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Tsai, Kai-Chen, Lo, Tien-Hong, Sun, Yun-Ting, Chen, Berlin
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866910235829993472
author Tsai, Kai-Chen
Lo, Tien-Hong
Sun, Yun-Ting
Chen, Berlin
author_facet Tsai, Kai-Chen
Lo, Tien-Hong
Sun, Yun-Ting
Chen, Berlin
contents Contextual biasing is essential to improving the recognition of rare and domain-specific words in an automatic speech recognition (ASR) system. While numerous methods have been proposed in recent years, most of them focus on offline settings and do not explicitly address the challenges of streaming ASR. For example, CTC-based word spotting (CTC-WS) have demonstrated strong performance by directly detecting keywords from CTC log-probabilities, but they are limited to offline processing and require access to the full utterance. In This work, we present a streaming extension of CTC-WS for real-time contextual biasing. Our method maintains active keyword paths across audio chunks using a stateful token passing algorithm, enabling the detection of keywords that span multiple chunks. To ensure low latency and stable output, we introduce an incremental commitment mechanism that only emits segments guaranteed not to be affected by future audio, while deferring uncertain regions. This method naturally integrates with streaming ASR pipelines and does not require modifications to the underlying acoustic model or additional training, making it practical for real-world deployment. Experimental results show that our method reduces overall WER and effectively improves keyword F-score, demonstrating its effectiveness for real-time ASR applications.
format Preprint
id arxiv_https___arxiv_org_abs_2605_18222
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Contextual Biasing for Streaming ASR via CTC-based Word Spotting
Tsai, Kai-Chen
Lo, Tien-Hong
Sun, Yun-Ting
Chen, Berlin
Audio and Speech Processing
Contextual biasing is essential to improving the recognition of rare and domain-specific words in an automatic speech recognition (ASR) system. While numerous methods have been proposed in recent years, most of them focus on offline settings and do not explicitly address the challenges of streaming ASR. For example, CTC-based word spotting (CTC-WS) have demonstrated strong performance by directly detecting keywords from CTC log-probabilities, but they are limited to offline processing and require access to the full utterance. In This work, we present a streaming extension of CTC-WS for real-time contextual biasing. Our method maintains active keyword paths across audio chunks using a stateful token passing algorithm, enabling the detection of keywords that span multiple chunks. To ensure low latency and stable output, we introduce an incremental commitment mechanism that only emits segments guaranteed not to be affected by future audio, while deferring uncertain regions. This method naturally integrates with streaming ASR pipelines and does not require modifications to the underlying acoustic model or additional training, making it practical for real-world deployment. Experimental results show that our method reduces overall WER and effectively improves keyword F-score, demonstrating its effectiveness for real-time ASR applications.
title Contextual Biasing for Streaming ASR via CTC-based Word Spotting
topic Audio and Speech Processing
url https://arxiv.org/abs/2605.18222