SSTFormer: Bridging Spiking Neural Network and Memory Support Transformer for Frame-Event based Recognition

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Wang, Xiao, Rong, Yao, Wu, Zongzhen, Zhu, Lin, Jiang, Bo, Tang, Jin, Tian, Yonghong
Formato: Preprint
Publicado: 2023
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912359035961344
author Wang, Xiao
Rong, Yao
Wu, Zongzhen
Zhu, Lin
Jiang, Bo
Tang, Jin
Tian, Yonghong
author_facet Wang, Xiao
Rong, Yao
Wu, Zongzhen
Zhu, Lin
Jiang, Bo
Tang, Jin
Tian, Yonghong
contents Event camera-based pattern recognition is a newly arising research topic in recent years. Current researchers usually transform the event streams into images, graphs, or voxels, and adopt deep neural networks for event-based classification. Although good performance can be achieved on simple event recognition datasets, however, their results may be still limited due to the following two issues. Firstly, they adopt spatial sparse event streams for recognition only, which may fail to capture the color and detailed texture information well. Secondly, they adopt either Spiking Neural Networks (SNN) for energy-efficient recognition with suboptimal results, or Artificial Neural Networks (ANN) for energy-intensive, high-performance recognition. However, seldom of them consider achieving a balance between these two aspects. In this paper, we formally propose to recognize patterns by fusing RGB frames and event streams simultaneously and propose a new RGB frame-event recognition framework to address the aforementioned issues. The proposed method contains four main modules, i.e., memory support Transformer network for RGB frame encoding, spiking neural network for raw event stream encoding, multi-modal bottleneck fusion module for RGB-Event feature aggregation, and prediction head. Due to the scarce of RGB-Event based classification dataset, we also propose a large-scale PokerEvent dataset which contains 114 classes, and 27102 frame-event pairs recorded using a DVS346 event camera. Extensive experiments on two RGB-Event based classification datasets fully validated the effectiveness of our proposed framework. We hope this work will boost the development of pattern recognition by fusing RGB frames and event streams. Both our dataset and source code of this work will be released at https://github.com/Event-AHU/SSTFormer
format Preprint
id arxiv_https___arxiv_org_abs_2308_04369
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle SSTFormer: Bridging Spiking Neural Network and Memory Support Transformer for Frame-Event based Recognition
Wang, Xiao
Rong, Yao
Wu, Zongzhen
Zhu, Lin
Jiang, Bo
Tang, Jin
Tian, Yonghong
Computer Vision and Pattern Recognition
Multimedia
Neural and Evolutionary Computing
Event camera-based pattern recognition is a newly arising research topic in recent years. Current researchers usually transform the event streams into images, graphs, or voxels, and adopt deep neural networks for event-based classification. Although good performance can be achieved on simple event recognition datasets, however, their results may be still limited due to the following two issues. Firstly, they adopt spatial sparse event streams for recognition only, which may fail to capture the color and detailed texture information well. Secondly, they adopt either Spiking Neural Networks (SNN) for energy-efficient recognition with suboptimal results, or Artificial Neural Networks (ANN) for energy-intensive, high-performance recognition. However, seldom of them consider achieving a balance between these two aspects. In this paper, we formally propose to recognize patterns by fusing RGB frames and event streams simultaneously and propose a new RGB frame-event recognition framework to address the aforementioned issues. The proposed method contains four main modules, i.e., memory support Transformer network for RGB frame encoding, spiking neural network for raw event stream encoding, multi-modal bottleneck fusion module for RGB-Event feature aggregation, and prediction head. Due to the scarce of RGB-Event based classification dataset, we also propose a large-scale PokerEvent dataset which contains 114 classes, and 27102 frame-event pairs recorded using a DVS346 event camera. Extensive experiments on two RGB-Event based classification datasets fully validated the effectiveness of our proposed framework. We hope this work will boost the development of pattern recognition by fusing RGB frames and event streams. Both our dataset and source code of this work will be released at https://github.com/Event-AHU/SSTFormer
title SSTFormer: Bridging Spiking Neural Network and Memory Support Transformer for Frame-Event based Recognition
topic Computer Vision and Pattern Recognition
Multimedia
Neural and Evolutionary Computing
url https://arxiv.org/abs/2308.04369