Saved in:
Bibliographic Details
Main Authors: Li, Dianze, Li, Jianing, Liu, Xu, Zhou, Zhaokun, Fan, Xiaopeng, Tian, Yonghong
Format: Preprint
Published: 2024
Subjects:
Online Access:https://arxiv.org/abs/2411.18658
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910719647154176
author Li, Dianze
Li, Jianing
Liu, Xu
Zhou, Zhaokun
Fan, Xiaopeng
Tian, Yonghong
author_facet Li, Dianze
Li, Jianing
Liu, Xu
Zhou, Zhaokun
Fan, Xiaopeng
Tian, Yonghong
contents Combining the complementary benefits of frames and events has been widely used for object detection in challenging scenarios. However, most object detection methods use two independent Artificial Neural Network (ANN) branches, limiting cross-modality information interaction across the two visual streams and encountering challenges in extracting temporal cues from event streams with low power consumption. To address these challenges, we propose HDI-Former, a Hybrid Dynamic Interaction ANN-SNN Transformer, marking the first trial to design a directly trained hybrid ANN-SNN architecture for high-accuracy and energy-efficient object detection using frames and events. Technically, we first present a novel semantic-enhanced self-attention mechanism that strengthens the correlation between image encoding tokens within the ANN Transformer branch for better performance. Then, we design a Spiking Swin Transformer branch to model temporal cues from event streams with low power consumption. Finally, we propose a bio-inspired dynamic interaction mechanism between ANN and SNN sub-networks for cross-modality information interaction. The results demonstrate that our HDI-Former outperforms eleven state-of-the-art methods and our four baselines by a large margin. Our SNN branch also shows comparable performance to the ANN with the same architecture while consuming 10.57$\times$ less energy on the DSEC-Detection dataset. Our open-source code is available in the supplementary material.
format Preprint
id arxiv_https___arxiv_org_abs_2411_18658
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle HDI-Former: Hybrid Dynamic Interaction ANN-SNN Transformer for Object Detection Using Frames and Events
Li, Dianze
Li, Jianing
Liu, Xu
Zhou, Zhaokun
Fan, Xiaopeng
Tian, Yonghong
Computer Vision and Pattern Recognition
Combining the complementary benefits of frames and events has been widely used for object detection in challenging scenarios. However, most object detection methods use two independent Artificial Neural Network (ANN) branches, limiting cross-modality information interaction across the two visual streams and encountering challenges in extracting temporal cues from event streams with low power consumption. To address these challenges, we propose HDI-Former, a Hybrid Dynamic Interaction ANN-SNN Transformer, marking the first trial to design a directly trained hybrid ANN-SNN architecture for high-accuracy and energy-efficient object detection using frames and events. Technically, we first present a novel semantic-enhanced self-attention mechanism that strengthens the correlation between image encoding tokens within the ANN Transformer branch for better performance. Then, we design a Spiking Swin Transformer branch to model temporal cues from event streams with low power consumption. Finally, we propose a bio-inspired dynamic interaction mechanism between ANN and SNN sub-networks for cross-modality information interaction. The results demonstrate that our HDI-Former outperforms eleven state-of-the-art methods and our four baselines by a large margin. Our SNN branch also shows comparable performance to the ANN with the same architecture while consuming 10.57$\times$ less energy on the DSEC-Detection dataset. Our open-source code is available in the supplementary material.
title HDI-Former: Hybrid Dynamic Interaction ANN-SNN Transformer for Object Detection Using Frames and Events
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2411.18658