Spiking Patches: Asynchronous, Sparse, and Efficient Tokens for Event Cameras

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Øhrstrøm, Christoffer Koo, Güldenring, Ronja, Nalpantidis, Lazaros
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914124387057664
author Øhrstrøm, Christoffer Koo
Güldenring, Ronja
Nalpantidis, Lazaros
author_facet Øhrstrøm, Christoffer Koo
Güldenring, Ronja
Nalpantidis, Lazaros
contents We propose tokenization of events and present a tokenizer, Spiking Patches, specifically designed for event cameras. Given a stream of asynchronous and spatially sparse events, our goal is to discover an event representation that preserves these properties. Prior works have represented events as frames or as voxels. However, while these representations yield high accuracy, both frames and voxels are synchronous and decrease the spatial sparsity. Spiking Patches gives the means to preserve the unique properties of event cameras and we show in our experiments that this comes without sacrificing accuracy. We evaluate our tokenizer using a GNN, PCN, and a Transformer on gesture recognition and object detection. Tokens from Spiking Patches yield inference times that are up to 3.4x faster than voxel-based tokens and up to 10.4x faster than frames. We achieve this while matching their accuracy and even surpassing in some cases with absolute improvements up to 3.8 for gesture recognition and up to 1.4 for object detection. Thus, tokenization constitutes a novel direction in event-based vision and marks a step towards methods that preserve the properties of event cameras.
format Preprint
id arxiv_https___arxiv_org_abs_2510_26614
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Spiking Patches: Asynchronous, Sparse, and Efficient Tokens for Event Cameras
Øhrstrøm, Christoffer Koo
Güldenring, Ronja
Nalpantidis, Lazaros
Computer Vision and Pattern Recognition
Robotics
We propose tokenization of events and present a tokenizer, Spiking Patches, specifically designed for event cameras. Given a stream of asynchronous and spatially sparse events, our goal is to discover an event representation that preserves these properties. Prior works have represented events as frames or as voxels. However, while these representations yield high accuracy, both frames and voxels are synchronous and decrease the spatial sparsity. Spiking Patches gives the means to preserve the unique properties of event cameras and we show in our experiments that this comes without sacrificing accuracy. We evaluate our tokenizer using a GNN, PCN, and a Transformer on gesture recognition and object detection. Tokens from Spiking Patches yield inference times that are up to 3.4x faster than voxel-based tokens and up to 10.4x faster than frames. We achieve this while matching their accuracy and even surpassing in some cases with absolute improvements up to 3.8 for gesture recognition and up to 1.4 for object detection. Thus, tokenization constitutes a novel direction in event-based vision and marks a step towards methods that preserve the properties of event cameras.
title Spiking Patches: Asynchronous, Sparse, and Efficient Tokens for Event Cameras
topic Computer Vision and Pattern Recognition
Robotics
url https://arxiv.org/abs/2510.26614