Improving vision-inspired keyword spotting using dynamic module skipping in streaming conformer encoder

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bittar, Alexandre, Dixon, Paul, Samragh, Mohammad, Nishu, Kumari, Naik, Devang
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929296883318784
author Bittar, Alexandre
Dixon, Paul
Samragh, Mohammad
Nishu, Kumari
Naik, Devang
author_facet Bittar, Alexandre
Dixon, Paul
Samragh, Mohammad
Nishu, Kumari
Naik, Devang
contents Using a vision-inspired keyword spotting framework, we propose an architecture with input-dependent dynamic depth capable of processing streaming audio. Specifically, we extend a conformer encoder with trainable binary gates that allow us to dynamically skip network modules according to the input audio. Our approach improves detection and localization accuracy on continuous speech using Librispeech top-1000 most frequent words while maintaining a small memory footprint. The inclusion of gates also reduces the average amount of processing without affecting the overall performance. These benefits are shown to be even more pronounced using the Google speech commands dataset placed over background noise where up to 97% of the processing is skipped on non-speech inputs, therefore making our method particularly interesting for an always-on keyword spotter.
format Preprint
id arxiv_https___arxiv_org_abs_2309_00140
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Improving vision-inspired keyword spotting using dynamic module skipping in streaming conformer encoder
Bittar, Alexandre
Dixon, Paul
Samragh, Mohammad
Nishu, Kumari
Naik, Devang
Sound
Computer Vision and Pattern Recognition
Machine Learning
Audio and Speech Processing
Using a vision-inspired keyword spotting framework, we propose an architecture with input-dependent dynamic depth capable of processing streaming audio. Specifically, we extend a conformer encoder with trainable binary gates that allow us to dynamically skip network modules according to the input audio. Our approach improves detection and localization accuracy on continuous speech using Librispeech top-1000 most frequent words while maintaining a small memory footprint. The inclusion of gates also reduces the average amount of processing without affecting the overall performance. These benefits are shown to be even more pronounced using the Google speech commands dataset placed over background noise where up to 97% of the processing is skipped on non-speech inputs, therefore making our method particularly interesting for an always-on keyword spotter.
title Improving vision-inspired keyword spotting using dynamic module skipping in streaming conformer encoder
topic Sound
Computer Vision and Pattern Recognition
Machine Learning
Audio and Speech Processing
url https://arxiv.org/abs/2309.00140