Global-Local Convolution with Spiking Neural Networks for Energy-efficient Keyword Spotting

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Shuai, Zhang, Dehao, Shi, Kexin, Wang, Yuchen, Wei, Wenjie, Wu, Jibin, Zhang, Malu
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916294258851840
author Wang, Shuai
Zhang, Dehao
Shi, Kexin
Wang, Yuchen
Wei, Wenjie
Wu, Jibin
Zhang, Malu
author_facet Wang, Shuai
Zhang, Dehao
Shi, Kexin
Wang, Yuchen
Wei, Wenjie
Wu, Jibin
Zhang, Malu
contents Thanks to Deep Neural Networks (DNNs), the accuracy of Keyword Spotting (KWS) has made substantial progress. However, as KWS systems are usually implemented on edge devices, energy efficiency becomes a critical requirement besides performance. Here, we take advantage of spiking neural networks' energy efficiency and propose an end-to-end lightweight KWS model. The model consists of two innovative modules: 1) Global-Local Spiking Convolution (GLSC) module and 2) Bottleneck-PLIF module. Compared to the hand-crafted feature extraction methods, the GLSC module achieves speech feature extraction that is sparser, more energy-efficient, and yields better performance. The Bottleneck-PLIF module further processes the signals from GLSC with the aim to achieve higher accuracy with fewer parameters. Extensive experiments are conducted on the Google Speech Commands Dataset (V1 and V2). The results show our method achieves competitive performance among SNN-based KWS models with fewer parameters.
format Preprint
id arxiv_https___arxiv_org_abs_2406_13179
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Global-Local Convolution with Spiking Neural Networks for Energy-efficient Keyword Spotting
Wang, Shuai
Zhang, Dehao
Shi, Kexin
Wang, Yuchen
Wei, Wenjie
Wu, Jibin
Zhang, Malu
Sound
Artificial Intelligence
Neural and Evolutionary Computing
Audio and Speech Processing
Thanks to Deep Neural Networks (DNNs), the accuracy of Keyword Spotting (KWS) has made substantial progress. However, as KWS systems are usually implemented on edge devices, energy efficiency becomes a critical requirement besides performance. Here, we take advantage of spiking neural networks' energy efficiency and propose an end-to-end lightweight KWS model. The model consists of two innovative modules: 1) Global-Local Spiking Convolution (GLSC) module and 2) Bottleneck-PLIF module. Compared to the hand-crafted feature extraction methods, the GLSC module achieves speech feature extraction that is sparser, more energy-efficient, and yields better performance. The Bottleneck-PLIF module further processes the signals from GLSC with the aim to achieve higher accuracy with fewer parameters. Extensive experiments are conducted on the Google Speech Commands Dataset (V1 and V2). The results show our method achieves competitive performance among SNN-based KWS models with fewer parameters.
title Global-Local Convolution with Spiking Neural Networks for Energy-efficient Keyword Spotting
topic Sound
Artificial Intelligence
Neural and Evolutionary Computing
Audio and Speech Processing
url https://arxiv.org/abs/2406.13179