Saved in:
Bibliographic Details
Main Authors: Chane, Camille Simon, Niebur, Ernst, Benosman, Ryad, Ieng, Sio-Hoi
Format: Preprint
Published: 2024
Subjects:
Online Access:https://arxiv.org/abs/2401.05030
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916086834790400
author Chane, Camille Simon
Niebur, Ernst
Benosman, Ryad
Ieng, Sio-Hoi
author_facet Chane, Camille Simon
Niebur, Ernst
Benosman, Ryad
Ieng, Sio-Hoi
contents Selective attention is an essential mechanism to filter sensory input and to select only its most important components, allowing the capacity-limited cognitive structures of the brain to process them in detail. The saliency map model, originally developed to understand the process of selective attention in the primate visual system, has also been extensively used in computer vision. Due to the wide-spread use of frame-based video, this is how dynamic input from non-stationary scenes is commonly implemented in saliency maps. However, the temporal structure of this input modality is very different from that of the primate visual system. Retinal input to the brain is massively parallel, local rather than frame-based, asynchronous rather than synchronous, and transmitted in the form of discrete events, neuronal action potentials (spikes). These features are captured by event-based cameras. We show that a computational saliency model can be obtained organically from such vision sensors, at minimal computational cost. We assess the performance of the model by comparing its predictions with the distribution of overt attention (fixations) of human observers, and we make available an event-based dataset that can be used as ground truth for future studies.
format Preprint
id arxiv_https___arxiv_org_abs_2401_05030
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle An event-based implementation of saliency-based visual attention for rapid scene analysis
Chane, Camille Simon
Niebur, Ernst
Benosman, Ryad
Ieng, Sio-Hoi
Image and Video Processing
Signal Processing
Selective attention is an essential mechanism to filter sensory input and to select only its most important components, allowing the capacity-limited cognitive structures of the brain to process them in detail. The saliency map model, originally developed to understand the process of selective attention in the primate visual system, has also been extensively used in computer vision. Due to the wide-spread use of frame-based video, this is how dynamic input from non-stationary scenes is commonly implemented in saliency maps. However, the temporal structure of this input modality is very different from that of the primate visual system. Retinal input to the brain is massively parallel, local rather than frame-based, asynchronous rather than synchronous, and transmitted in the form of discrete events, neuronal action potentials (spikes). These features are captured by event-based cameras. We show that a computational saliency model can be obtained organically from such vision sensors, at minimal computational cost. We assess the performance of the model by comparing its predictions with the distribution of overt attention (fixations) of human observers, and we make available an event-based dataset that can be used as ground truth for future studies.
title An event-based implementation of saliency-based visual attention for rapid scene analysis
topic Image and Video Processing
Signal Processing
url https://arxiv.org/abs/2401.05030