SpikCommander: A High-performance Spiking Transformer with Multi-view Learning for Efficient Speech Command Recognition
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Jiaqi, Yu, Liutao, Shen, Xiongri, Guo, Sihang, Zhou, Chenlin, Zhao, Leilei, Zhong, Yi, Zhang, Zhiguo, Ma, Zhengyu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Efficient Speech Command Recognition Leveraging Spiking Neural Network and Curriculum Learning-based Knowledge Distillation
von: Wang, Jiaqi, et al.
Veröffentlicht: (2024)
von: Wang, Jiaqi, et al.
Veröffentlicht: (2024)
BiSpikCLM: A Spiking Language Model integrating Softmax-Free Spiking Attention and Spike-Aware Alignment Distillation
von: Guo, Sihang, et al.
Veröffentlicht: (2026)
von: Guo, Sihang, et al.
Veröffentlicht: (2026)
Moonshine: Speech Recognition for Live Transcription and Voice Commands
von: Jeffries, Nat, et al.
Veröffentlicht: (2024)
von: Jeffries, Nat, et al.
Veröffentlicht: (2024)
Hello Afrika: Speech Commands in Kinyarwanda
von: Igwegbe, George, et al.
Veröffentlicht: (2025)
von: Igwegbe, George, et al.
Veröffentlicht: (2025)
Speech Command Recognition Using LogNNet Reservoir Computing for Embedded Systems
von: Izotov, Yuriy, et al.
Veröffentlicht: (2025)
von: Izotov, Yuriy, et al.
Veröffentlicht: (2025)
S$^2$M-Former: Spiking Symmetric Mixing Branchformer for Brain Auditory Attention Detection
von: Wang, Jiaqi, et al.
Veröffentlicht: (2025)
von: Wang, Jiaqi, et al.
Veröffentlicht: (2025)
SynTTS-Commands: A Public Dataset for On-Device KWS via TTS-Synthesized Multilingual Speech
von: Gan, Lu, et al.
Veröffentlicht: (2025)
von: Gan, Lu, et al.
Veröffentlicht: (2025)
Improving Pretrained YAMNet for Enhanced Speech Command Detection via Transfer Learning
von: Lachenani, Sidahmed, et al.
Veröffentlicht: (2025)
von: Lachenani, Sidahmed, et al.
Veröffentlicht: (2025)
SVFormer: A Direct Training Spiking Transformer for Efficient Video Action Recognition
von: Yu, Liutao, et al.
Veröffentlicht: (2024)
von: Yu, Liutao, et al.
Veröffentlicht: (2024)
Advancing Airport Tower Command Recognition: Integrating Squeeze-and-Excitation and Broadcasted Residual Learning
von: Lin, Yuanxi, et al.
Veröffentlicht: (2024)
von: Lin, Yuanxi, et al.
Veröffentlicht: (2024)
Scalable Offline ASR for Command-Style Dictation in Courtrooms
von: Nethil, Kumarmanas, et al.
Veröffentlicht: (2025)
von: Nethil, Kumarmanas, et al.
Veröffentlicht: (2025)
Evaluating Synthetic Command Attacks on Smart Voice Assistants
von: He, Zhengxian, et al.
Veröffentlicht: (2024)
von: He, Zhengxian, et al.
Veröffentlicht: (2024)
Spiking and Event-driven Neuromorphic Mamba Models for Efficient Speech Recognition
von: Ahmed, Tauseef, et al.
Veröffentlicht: (2026)
von: Ahmed, Tauseef, et al.
Veröffentlicht: (2026)
Enhancing Synthetic Training Data for Speech Commands: From ASR-Based Filtering to Domain Adaptation in SSL Latent Space
von: Quintas, Sebastião, et al.
Veröffentlicht: (2024)
von: Quintas, Sebastião, et al.
Veröffentlicht: (2024)
PTS-SNN: A Prompt-Tuned Temporal Shift Spiking Neural Networks for Efficient Speech Emotion Recognition
von: Su, Xun, et al.
Veröffentlicht: (2026)
von: Su, Xun, et al.
Veröffentlicht: (2026)
Analyzing Multimodal Features of Spontaneous Voice Assistant Commands for Mild Cognitive Impairment Detection
von: Lin, Nana, et al.
Veröffentlicht: (2024)
von: Lin, Nana, et al.
Veröffentlicht: (2024)
Winner-Take-All Spiking Transformer for Language Modeling
von: Zhou, Chenlin, et al.
Veröffentlicht: (2026)
von: Zhou, Chenlin, et al.
Veröffentlicht: (2026)
Temporal-Frequency State Space Duality: An Efficient Paradigm for Speech Emotion Recognition
von: Zhao, Jiaqi, et al.
Veröffentlicht: (2024)
von: Zhao, Jiaqi, et al.
Veröffentlicht: (2024)
Efficient Finetuning for Dimensional Speech Emotion Recognition in the Age of Transformers
von: Sampath, Aneesha, et al.
Veröffentlicht: (2025)
von: Sampath, Aneesha, et al.
Veröffentlicht: (2025)
Breaking Through the Spike: Spike Window Decoding for Accelerated and Precise Automatic Speech Recognition
von: Zhang, Wei, et al.
Veröffentlicht: (2025)
von: Zhang, Wei, et al.
Veröffentlicht: (2025)
Evaluating Voice Command Pipelines for Drone Control: From STT and LLM to Direct Classification and Siamese Networks
von: Simões, Lucca Emmanuel Pineli, et al.
Veröffentlicht: (2024)
von: Simões, Lucca Emmanuel Pineli, et al.
Veröffentlicht: (2024)
DiTSE: High-Fidelity Generative Speech Enhancement via Latent Diffusion Transformers
von: Guimarães, Heitor R., et al.
Veröffentlicht: (2025)
von: Guimarães, Heitor R., et al.
Veröffentlicht: (2025)
Adaptive Spiking Neurons for Vision and Language Modeling
von: Zhou, Chenlin, et al.
Veröffentlicht: (2026)
von: Zhou, Chenlin, et al.
Veröffentlicht: (2026)
Efficient Long-Form Speech Recognition for General Speech In-Context Learning
von: Yen, Hao, et al.
Veröffentlicht: (2024)
von: Yen, Hao, et al.
Veröffentlicht: (2024)
An Efficient Transfer Learning Method Based on Adapter with Local Attributes for Speech Emotion Recognition
von: Song, Haoyu, et al.
Veröffentlicht: (2025)
von: Song, Haoyu, et al.
Veröffentlicht: (2025)
SimpleSpeech: Towards Simple and Efficient Text-to-Speech with Scalar Latent Transformer Diffusion Models
von: Yang, Dongchao, et al.
Veröffentlicht: (2024)
von: Yang, Dongchao, et al.
Veröffentlicht: (2024)
Explainable Transformer-CNN Fusion for Noise-Robust Speech Emotion Recognition
von: Chakrabarty, Sudip, et al.
Veröffentlicht: (2025)
von: Chakrabarty, Sudip, et al.
Veröffentlicht: (2025)
SpikeVox: Towards Energy-Efficient Speech Therapy Framework with Spike-driven Generative Language Models
von: Putra, Rachmad Vidya Wicaksana, et al.
Veröffentlicht: (2025)
von: Putra, Rachmad Vidya Wicaksana, et al.
Veröffentlicht: (2025)
Spikingformer: A Key Foundation Model for Spiking Neural Networks
von: Zhou, Chenlin, et al.
Veröffentlicht: (2023)
von: Zhou, Chenlin, et al.
Veröffentlicht: (2023)
ROSE: A Recognition-Oriented Speech Enhancement Framework in Air Traffic Control Using Multi-Objective Learning
von: Yu, Xincheng, et al.
Veröffentlicht: (2023)
von: Yu, Xincheng, et al.
Veröffentlicht: (2023)
Multi-Channel Speech Enhancement for Cocktail Party Speech Emotion Recognition
von: Chen, Youjun, et al.
Veröffentlicht: (2026)
von: Chen, Youjun, et al.
Veröffentlicht: (2026)
Speech Recognition Transformers: Topological-lingualism Perspective
von: Singh, Shruti, et al.
Veröffentlicht: (2024)
von: Singh, Shruti, et al.
Veröffentlicht: (2024)
A Transcription Prompt-based Efficient Audio Large Language Model for Robust Speech Recognition
von: Li, Yangze, et al.
Veröffentlicht: (2024)
von: Li, Yangze, et al.
Veröffentlicht: (2024)
SyncSpeech: Efficient and Low-Latency Text-to-Speech based on Temporal Masked Transformer
von: Sheng, Zhengyan, et al.
Veröffentlicht: (2025)
von: Sheng, Zhengyan, et al.
Veröffentlicht: (2025)
GigaAM: Efficient Self-Supervised Learner for Speech Recognition
von: Kutsakov, Aleksandr, et al.
Veröffentlicht: (2025)
von: Kutsakov, Aleksandr, et al.
Veröffentlicht: (2025)
Towards Ultra-Low-Power Neuromorphic Speech Enhancement with Spiking-FullSubNet
von: Hao, Xiang, et al.
Veröffentlicht: (2024)
von: Hao, Xiang, et al.
Veröffentlicht: (2024)
Speech Emotion Recognition via Entropy-Aware Score Selection
von: Chua, ChenYi, et al.
Veröffentlicht: (2025)
von: Chua, ChenYi, et al.
Veröffentlicht: (2025)
Toward Fairness in Speech Recognition: Discovery and mitigation of performance disparities
von: Dheram, Pranav, et al.
Veröffentlicht: (2022)
von: Dheram, Pranav, et al.
Veröffentlicht: (2022)
Hybrid CNN-Transformer Architecture for Arabic Speech Emotion Recognition
von: Gheffari, Youcef Soufiane, et al.
Veröffentlicht: (2026)
von: Gheffari, Youcef Soufiane, et al.
Veröffentlicht: (2026)
Computation and Parameter Efficient Multi-Modal Fusion Transformer for Cued Speech Recognition
von: Liu, Lei, et al.
Veröffentlicht: (2024)
von: Liu, Lei, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Efficient Speech Command Recognition Leveraging Spiking Neural Network and Curriculum Learning-based Knowledge Distillation
von: Wang, Jiaqi, et al.
Veröffentlicht: (2024) -
BiSpikCLM: A Spiking Language Model integrating Softmax-Free Spiking Attention and Spike-Aware Alignment Distillation
von: Guo, Sihang, et al.
Veröffentlicht: (2026) -
Moonshine: Speech Recognition for Live Transcription and Voice Commands
von: Jeffries, Nat, et al.
Veröffentlicht: (2024) -
Hello Afrika: Speech Commands in Kinyarwanda
von: Igwegbe, George, et al.
Veröffentlicht: (2025) -
Speech Command Recognition Using LogNNet Reservoir Computing for Embedded Systems
von: Izotov, Yuriy, et al.
Veröffentlicht: (2025)