Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Hao, Haiqing, Sui, Zhipeng, Zou, Rong, Dai, Zijia, Zubić, Nikola, Scaramuzza, Davide, Wang, Wenhui
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:https://arxiv.org/abs/2603.06228
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866918375646560256
author Hao, Haiqing
Sui, Zhipeng
Zou, Rong
Dai, Zijia
Zubić, Nikola
Scaramuzza, Davide
Wang, Wenhui
author_facet Hao, Haiqing
Sui, Zhipeng
Zou, Rong
Dai, Zijia
Zubić, Nikola
Scaramuzza, Davide
Wang, Wenhui
contents Event cameras provide sequential visual data with spatial sparsity and high temporal resolution, making them attractive for low-latency object detection. Existing asynchronous event-based neural networks realize this low-latency advantage by updating predictions event-by-event, but still suffer from two bottlenecks: recurrent architectures are difficult to train efficiently on long sequences, and improving accuracy often increases per-event computation and latency. Linear attention is appealing in this setting because it supports parallel training and recurrent inference. However, standard linear attention updates a global state for every event, yielding a poor accuracy-efficiency trade-off, which is problematic for object detection, where fine-grained representations and thus states are preferred. The key challenge is therefore to introduce sparse state activation that exploits event sparsity while preserving efficient parallel training. We propose Spatially-Sparse Linear Attention (SSLA), which introduces a mixture-of-spaces state decomposition and a scatter-compute-gather training procedure, enabling state-level sparsity as well as training parallelism. Built on SSLA, we develop an end-to-end asynchronous linear attention model, SSLA-Det, for event-based object detection. On Gen1 and N-Caltech101, SSLA-Det achieves state-of-the-art accuracy among asynchronous methods, reaching 0.375 mAP and 0.515 mAP, respectively, while reducing per-event computation by more than 20 times compared to the strongest prior asynchronous baseline, demonstrating the potential of linear attention for low-latency event-based vision.
format Preprint
id arxiv_https___arxiv_org_abs_2603_06228
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Low-latency Event-based Object Detection with Spatially-Sparse Linear Attention
Hao, Haiqing
Sui, Zhipeng
Zou, Rong
Dai, Zijia
Zubić, Nikola
Scaramuzza, Davide
Wang, Wenhui
Computer Vision and Pattern Recognition
Event cameras provide sequential visual data with spatial sparsity and high temporal resolution, making them attractive for low-latency object detection. Existing asynchronous event-based neural networks realize this low-latency advantage by updating predictions event-by-event, but still suffer from two bottlenecks: recurrent architectures are difficult to train efficiently on long sequences, and improving accuracy often increases per-event computation and latency. Linear attention is appealing in this setting because it supports parallel training and recurrent inference. However, standard linear attention updates a global state for every event, yielding a poor accuracy-efficiency trade-off, which is problematic for object detection, where fine-grained representations and thus states are preferred. The key challenge is therefore to introduce sparse state activation that exploits event sparsity while preserving efficient parallel training. We propose Spatially-Sparse Linear Attention (SSLA), which introduces a mixture-of-spaces state decomposition and a scatter-compute-gather training procedure, enabling state-level sparsity as well as training parallelism. Built on SSLA, we develop an end-to-end asynchronous linear attention model, SSLA-Det, for event-based object detection. On Gen1 and N-Caltech101, SSLA-Det achieves state-of-the-art accuracy among asynchronous methods, reaching 0.375 mAP and 0.515 mAP, respectively, while reducing per-event computation by more than 20 times compared to the strongest prior asynchronous baseline, demonstrating the potential of linear attention for low-latency event-based vision.
title Low-latency Event-based Object Detection with Spatially-Sparse Linear Attention
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2603.06228