Winner-Take-All Spiking Transformer for Language Modeling

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Zhou, Chenlin, Guo, Sihang, Wang, Jiaqi, Ma, Dongyang, Che, Kaiwei, Chen, Baiyu, Meng, Qingyan, Ma, Zhengyu, Tian, Yonghong
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866911588523442176
author Zhou, Chenlin
Guo, Sihang
Wang, Jiaqi
Ma, Dongyang
Che, Kaiwei
Chen, Baiyu
Meng, Qingyan
Ma, Zhengyu
Tian, Yonghong
author_facet Zhou, Chenlin
Guo, Sihang
Wang, Jiaqi
Ma, Dongyang
Che, Kaiwei
Chen, Baiyu
Meng, Qingyan
Ma, Zhengyu
Tian, Yonghong
contents Spiking Transformers, which combine the scalability of Transformers with the sparse, energy-efficient property of Spiking Neural Networks (SNNs), have achieved impressive results in neuromorphic and vision tasks and attracted increasing attention. However, existing directly trained spiking transformers primarily focus on vision tasks. For language modeling with spiking transformer, convergence relies heavily on softmax-based spiking self-attention, which incurs high energy costs and poses challenges for neuromorphic deployment. To address this issue, we introduce Winner-Take-All (WTA) mechanisms into spiking transformers and propose two novel softmax-free, spike-driven self-attention modules: WTA Spiking Self-Attention (WSSA) and Causal WTA Spiking Self-Attention (CWSSA). Based on them, we design WTA-based Encoder-only Spiking Transformer (WE-Spikingformer) for masked language modeling and WTA-based Decoder-only Spiking Transformer (WD-Spikingformer) for causal language modeling, systematically exploring softmax-free, spiking-driven Transformer architectures trained end-to-end for natural language processing tasks. Extensive experiments on 16 datasets spanning natural language understanding, question-answering tasks, and commonsense reasoning tasks validate the effectiveness of our approach and highlight the promise of spiking transformers for general language modeling and energy-efficient artificial intelligence.
format Preprint
id arxiv_https___arxiv_org_abs_2604_11321
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Winner-Take-All Spiking Transformer for Language Modeling
Zhou, Chenlin
Guo, Sihang
Wang, Jiaqi
Ma, Dongyang
Che, Kaiwei
Chen, Baiyu
Meng, Qingyan
Ma, Zhengyu
Tian, Yonghong
Neural and Evolutionary Computing
Spiking Transformers, which combine the scalability of Transformers with the sparse, energy-efficient property of Spiking Neural Networks (SNNs), have achieved impressive results in neuromorphic and vision tasks and attracted increasing attention. However, existing directly trained spiking transformers primarily focus on vision tasks. For language modeling with spiking transformer, convergence relies heavily on softmax-based spiking self-attention, which incurs high energy costs and poses challenges for neuromorphic deployment. To address this issue, we introduce Winner-Take-All (WTA) mechanisms into spiking transformers and propose two novel softmax-free, spike-driven self-attention modules: WTA Spiking Self-Attention (WSSA) and Causal WTA Spiking Self-Attention (CWSSA). Based on them, we design WTA-based Encoder-only Spiking Transformer (WE-Spikingformer) for masked language modeling and WTA-based Decoder-only Spiking Transformer (WD-Spikingformer) for causal language modeling, systematically exploring softmax-free, spiking-driven Transformer architectures trained end-to-end for natural language processing tasks. Extensive experiments on 16 datasets spanning natural language understanding, question-answering tasks, and commonsense reasoning tasks validate the effectiveness of our approach and highlight the promise of spiking transformers for general language modeling and energy-efficient artificial intelligence.
title Winner-Take-All Spiking Transformer for Language Modeling
topic Neural and Evolutionary Computing
url https://arxiv.org/abs/2604.11321