SpikCommander: A High-performance Spiking Transformer with Multi-view Learning for Efficient Speech Command Recognition

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Wang, Jiaqi, Yu, Liutao, Shen, Xiongri, Guo, Sihang, Zhou, Chenlin, Zhao, Leilei, Zhong, Yi, Zhang, Zhiguo, Ma, Zhengyu
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866908773955665920
author Wang, Jiaqi
Yu, Liutao
Shen, Xiongri
Guo, Sihang
Zhou, Chenlin
Zhao, Leilei
Zhong, Yi
Zhang, Zhiguo
Ma, Zhengyu
author_facet Wang, Jiaqi
Yu, Liutao
Shen, Xiongri
Guo, Sihang
Zhou, Chenlin
Zhao, Leilei
Zhong, Yi
Zhang, Zhiguo
Ma, Zhengyu
contents Spiking neural networks (SNNs) offer a promising path toward energy-efficient speech command recognition (SCR) by leveraging their event-driven processing paradigm. However, existing SNN-based SCR methods often struggle to capture rich temporal dependencies and contextual information from speech due to limited temporal modeling and binary spike-based representations. To address these challenges, we first introduce the multi-view spiking temporal-aware self-attention (MSTASA) module, which combines effective spiking temporal-aware attention with a multi-view learning framework to model complementary temporal dependencies in speech commands. Building on MSTASA, we further propose SpikCommander, a fully spike-driven transformer architecture that integrates MSTASA with a spiking contextual refinement channel MLP (SCR-MLP) to jointly enhance temporal context modeling and channel-wise feature integration. We evaluate our method on three benchmark datasets: the Spiking Heidelberg Dataset (SHD), the Spiking Speech Commands (SSC), and the Google Speech Commands V2 (GSC). Extensive experiments demonstrate that SpikCommander consistently outperforms state-of-the-art (SOTA) SNN approaches with fewer parameters under comparable time steps, highlighting its effectiveness and efficiency for robust speech command recognition.
format Preprint
id arxiv_https___arxiv_org_abs_2511_07883
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SpikCommander: A High-performance Spiking Transformer with Multi-view Learning for Efficient Speech Command Recognition
Wang, Jiaqi
Yu, Liutao
Shen, Xiongri
Guo, Sihang
Zhou, Chenlin
Zhao, Leilei
Zhong, Yi
Zhang, Zhiguo
Ma, Zhengyu
Sound
Machine Learning
Spiking neural networks (SNNs) offer a promising path toward energy-efficient speech command recognition (SCR) by leveraging their event-driven processing paradigm. However, existing SNN-based SCR methods often struggle to capture rich temporal dependencies and contextual information from speech due to limited temporal modeling and binary spike-based representations. To address these challenges, we first introduce the multi-view spiking temporal-aware self-attention (MSTASA) module, which combines effective spiking temporal-aware attention with a multi-view learning framework to model complementary temporal dependencies in speech commands. Building on MSTASA, we further propose SpikCommander, a fully spike-driven transformer architecture that integrates MSTASA with a spiking contextual refinement channel MLP (SCR-MLP) to jointly enhance temporal context modeling and channel-wise feature integration. We evaluate our method on three benchmark datasets: the Spiking Heidelberg Dataset (SHD), the Spiking Speech Commands (SSC), and the Google Speech Commands V2 (GSC). Extensive experiments demonstrate that SpikCommander consistently outperforms state-of-the-art (SOTA) SNN approaches with fewer parameters under comparable time steps, highlighting its effectiveness and efficiency for robust speech command recognition.
title SpikCommander: A High-performance Spiking Transformer with Multi-view Learning for Efficient Speech Command Recognition
topic Sound
Machine Learning
url https://arxiv.org/abs/2511.07883