Spike2Former: Efficient Spiking Transformer for High-performance Image Segmentation

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Lei, Zhenxin, Yao, Man, Hu, Jiakui, Luo, Xinhao, Lu, Yanye, Xu, Bo, Li, Guoqi
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866909434023772160
author Lei, Zhenxin
Yao, Man
Hu, Jiakui
Luo, Xinhao
Lu, Yanye
Xu, Bo
Li, Guoqi
author_facet Lei, Zhenxin
Yao, Man
Hu, Jiakui
Luo, Xinhao
Lu, Yanye
Xu, Bo
Li, Guoqi
contents Spiking Neural Networks (SNNs) have a low-power advantage but perform poorly in image segmentation tasks. The reason is that directly converting neural networks with complex architectural designs for segmentation tasks into spiking versions leads to performance degradation and non-convergence. To address this challenge, we first identify the modules in the architecture design that lead to the severe reduction in spike firing, make targeted improvements, and propose Spike2Former architecture. Second, we propose normalized integer spiking neurons to solve the training stability problem of SNNs with complex architectures. We set a new state-of-the-art for SNNs in various semantic segmentation datasets, with a significant improvement of +12.7% mIoU and 5.0 efficiency on ADE20K, +14.3% mIoU and 5.2 efficiency on VOC2012, and +9.1% mIoU and 6.6 efficiency on CityScapes.
format Preprint
id arxiv_https___arxiv_org_abs_2412_14587
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Spike2Former: Efficient Spiking Transformer for High-performance Image Segmentation
Lei, Zhenxin
Yao, Man
Hu, Jiakui
Luo, Xinhao
Lu, Yanye
Xu, Bo
Li, Guoqi
Computer Vision and Pattern Recognition
Artificial Intelligence
Neural and Evolutionary Computing
Spiking Neural Networks (SNNs) have a low-power advantage but perform poorly in image segmentation tasks. The reason is that directly converting neural networks with complex architectural designs for segmentation tasks into spiking versions leads to performance degradation and non-convergence. To address this challenge, we first identify the modules in the architecture design that lead to the severe reduction in spike firing, make targeted improvements, and propose Spike2Former architecture. Second, we propose normalized integer spiking neurons to solve the training stability problem of SNNs with complex architectures. We set a new state-of-the-art for SNNs in various semantic segmentation datasets, with a significant improvement of +12.7% mIoU and 5.0 efficiency on ADE20K, +14.3% mIoU and 5.2 efficiency on VOC2012, and +9.1% mIoU and 6.6 efficiency on CityScapes.
title Spike2Former: Efficient Spiking Transformer for High-performance Image Segmentation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Neural and Evolutionary Computing
url https://arxiv.org/abs/2412.14587