Frequency & Channel Attention Network for Small Footprint Noisy Spoken Keyword Spotting

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Lin, Yuanxi, Gapanyuk, Yuriy Evgenyevich
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914891330224128
author Lin, Yuanxi
Gapanyuk, Yuriy Evgenyevich
author_facet Lin, Yuanxi
Gapanyuk, Yuriy Evgenyevich
contents In this paper, we aim to improve the robustness of Keyword Spotting (KWS) systems in noisy environments while keeping a small memory footprint. We propose a new convolutional neural network (CNN) called FCA-Net, which combines mixer unit-based feature interaction with a two-dimensional convolution-based attention module. First, we introduce and compare lightweight attention methods to enhance noise robustness in CNN. Then, we propose an attention module that creates fine-grained attention weights to capture channel and frequency-specific information, boosting the model's ability to handle noisy conditions. By combining the mixer unit-based feature interaction with the attention module, we enhance performance. Additionally, we use a curriculum-based multi-condition training strategy. Our experiments show that our system outperforms current state-of-the-art solutions for small-footprint KWS in noisy environments, making it reliable for real-world use.
format Preprint
id arxiv_https___arxiv_org_abs_2407_19834
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Frequency & Channel Attention Network for Small Footprint Noisy Spoken Keyword Spotting
Lin, Yuanxi
Gapanyuk, Yuriy Evgenyevich
Audio and Speech Processing
Sound
In this paper, we aim to improve the robustness of Keyword Spotting (KWS) systems in noisy environments while keeping a small memory footprint. We propose a new convolutional neural network (CNN) called FCA-Net, which combines mixer unit-based feature interaction with a two-dimensional convolution-based attention module. First, we introduce and compare lightweight attention methods to enhance noise robustness in CNN. Then, we propose an attention module that creates fine-grained attention weights to capture channel and frequency-specific information, boosting the model's ability to handle noisy conditions. By combining the mixer unit-based feature interaction with the attention module, we enhance performance. Additionally, we use a curriculum-based multi-condition training strategy. Our experiments show that our system outperforms current state-of-the-art solutions for small-footprint KWS in noisy environments, making it reliable for real-world use.
title Frequency & Channel Attention Network for Small Footprint Noisy Spoken Keyword Spotting
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2407.19834