Spike2Former: Efficient Spiking Transformer for High-performance Image Segmentation
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866909434023772160 |
|---|---|
| author | Lei, Zhenxin Yao, Man Hu, Jiakui Luo, Xinhao Lu, Yanye Xu, Bo Li, Guoqi |
| author_facet | Lei, Zhenxin Yao, Man Hu, Jiakui Luo, Xinhao Lu, Yanye Xu, Bo Li, Guoqi |
| contents | Spiking Neural Networks (SNNs) have a low-power advantage but perform poorly in image segmentation tasks. The reason is that directly converting neural networks with complex architectural designs for segmentation tasks into spiking versions leads to performance degradation and non-convergence. To address this challenge, we first identify the modules in the architecture design that lead to the severe reduction in spike firing, make targeted improvements, and propose Spike2Former architecture. Second, we propose normalized integer spiking neurons to solve the training stability problem of SNNs with complex architectures. We set a new state-of-the-art for SNNs in various semantic segmentation datasets, with a significant improvement of +12.7% mIoU and 5.0 efficiency on ADE20K, +14.3% mIoU and 5.2 efficiency on VOC2012, and +9.1% mIoU and 6.6 efficiency on CityScapes. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2412_14587 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Spike2Former: Efficient Spiking Transformer for High-performance Image Segmentation Lei, Zhenxin Yao, Man Hu, Jiakui Luo, Xinhao Lu, Yanye Xu, Bo Li, Guoqi Computer Vision and Pattern Recognition Artificial Intelligence Neural and Evolutionary Computing Spiking Neural Networks (SNNs) have a low-power advantage but perform poorly in image segmentation tasks. The reason is that directly converting neural networks with complex architectural designs for segmentation tasks into spiking versions leads to performance degradation and non-convergence. To address this challenge, we first identify the modules in the architecture design that lead to the severe reduction in spike firing, make targeted improvements, and propose Spike2Former architecture. Second, we propose normalized integer spiking neurons to solve the training stability problem of SNNs with complex architectures. We set a new state-of-the-art for SNNs in various semantic segmentation datasets, with a significant improvement of +12.7% mIoU and 5.0 efficiency on ADE20K, +14.3% mIoU and 5.2 efficiency on VOC2012, and +9.1% mIoU and 6.6 efficiency on CityScapes. |
| title | Spike2Former: Efficient Spiking Transformer for High-performance Image Segmentation |
| topic | Computer Vision and Pattern Recognition Artificial Intelligence Neural and Evolutionary Computing |
| url | https://arxiv.org/abs/2412.14587 |