Streaming Keyword Spotting Boosted by Cross-layer Discrimination Consistency

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Xi, Yu, Li, Haoyu, Gu, Xiaoyu, Li, Hao, Jiang, Yidi, Yu, Kai
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915077999820800
author Xi, Yu
Li, Haoyu
Gu, Xiaoyu
Li, Hao
Jiang, Yidi
Yu, Kai
author_facet Xi, Yu
Li, Haoyu
Gu, Xiaoyu
Li, Hao
Jiang, Yidi
Yu, Kai
contents Connectionist Temporal Classification (CTC), a non-autoregressive training criterion, is widely used in online keyword spotting (KWS). However, existing CTC-based KWS decoding strategies either rely on Automatic Speech Recognition (ASR), which performs suboptimally due to its broad search over the acoustic space without keyword-specific optimization, or on KWS-specific decoding graphs, which are complex to implement and maintain. In this work, we propose a streaming decoding algorithm enhanced by Cross-layer Discrimination Consistency (CDC), tailored for CTC-based KWS. Specifically, we introduce a streamlined yet effective decoding algorithm capable of detecting the start of the keyword at any arbitrary position. Furthermore, we leverage discrimination consistency information across layers to better differentiate between positive and false alarm samples. Our experiments on both clean and noisy Hey Snips datasets show that the proposed streaming decoding strategy outperforms ASR-based and graph-based KWS baselines. The CDC-boosted decoding further improves performance, yielding an average absolute recall improvement of 6.8% and a 46.3% relative reduction in the miss rate compared to the graph-based KWS baseline, with a very low false alarm rate of 0.05 per hour.
format Preprint
id arxiv_https___arxiv_org_abs_2412_12635
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Streaming Keyword Spotting Boosted by Cross-layer Discrimination Consistency
Xi, Yu
Li, Haoyu
Gu, Xiaoyu
Li, Hao
Jiang, Yidi
Yu, Kai
Audio and Speech Processing
Sound
Connectionist Temporal Classification (CTC), a non-autoregressive training criterion, is widely used in online keyword spotting (KWS). However, existing CTC-based KWS decoding strategies either rely on Automatic Speech Recognition (ASR), which performs suboptimally due to its broad search over the acoustic space without keyword-specific optimization, or on KWS-specific decoding graphs, which are complex to implement and maintain. In this work, we propose a streaming decoding algorithm enhanced by Cross-layer Discrimination Consistency (CDC), tailored for CTC-based KWS. Specifically, we introduce a streamlined yet effective decoding algorithm capable of detecting the start of the keyword at any arbitrary position. Furthermore, we leverage discrimination consistency information across layers to better differentiate between positive and false alarm samples. Our experiments on both clean and noisy Hey Snips datasets show that the proposed streaming decoding strategy outperforms ASR-based and graph-based KWS baselines. The CDC-boosted decoding further improves performance, yielding an average absolute recall improvement of 6.8% and a 46.3% relative reduction in the miss rate compared to the graph-based KWS baseline, with a very low false alarm rate of 0.05 per hour.
title Streaming Keyword Spotting Boosted by Cross-layer Discrimination Consistency
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2412.12635