EV-LayerSegNet: Self-supervised Motion Segmentation using Event Cameras

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Farah, Youssef, Paredes-Vallés, Federico, De Croon, Guido, Humais, Muhammad Ahmed, Sajwani, Hussain, Zweiri, Yahya
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908397715062784
author Farah, Youssef
Paredes-Vallés, Federico
De Croon, Guido
Humais, Muhammad Ahmed
Sajwani, Hussain
Zweiri, Yahya
author_facet Farah, Youssef
Paredes-Vallés, Federico
De Croon, Guido
Humais, Muhammad Ahmed
Sajwani, Hussain
Zweiri, Yahya
contents Event cameras are novel bio-inspired sensors that capture motion dynamics with much higher temporal resolution than traditional cameras, since pixels react asynchronously to brightness changes. They are therefore better suited for tasks involving motion such as motion segmentation. However, training event-based networks still represents a difficult challenge, as obtaining ground truth is very expensive, error-prone and limited in frequency. In this article, we introduce EV-LayerSegNet, a self-supervised CNN for event-based motion segmentation. Inspired by a layered representation of the scene dynamics, we show that it is possible to learn affine optical flow and segmentation masks separately, and use them to deblur the input events. The deblurring quality is then measured and used as self-supervised learning loss. We train and test the network on a simulated dataset with only affine motion, achieving IoU and detection rate up to 71% and 87% respectively.
format Preprint
id arxiv_https___arxiv_org_abs_2506_06596
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle EV-LayerSegNet: Self-supervised Motion Segmentation using Event Cameras
Farah, Youssef
Paredes-Vallés, Federico
De Croon, Guido
Humais, Muhammad Ahmed
Sajwani, Hussain
Zweiri, Yahya
Computer Vision and Pattern Recognition
Event cameras are novel bio-inspired sensors that capture motion dynamics with much higher temporal resolution than traditional cameras, since pixels react asynchronously to brightness changes. They are therefore better suited for tasks involving motion such as motion segmentation. However, training event-based networks still represents a difficult challenge, as obtaining ground truth is very expensive, error-prone and limited in frequency. In this article, we introduce EV-LayerSegNet, a self-supervised CNN for event-based motion segmentation. Inspired by a layered representation of the scene dynamics, we show that it is possible to learn affine optical flow and segmentation masks separately, and use them to deblur the input events. The deblurring quality is then measured and used as self-supervised learning loss. We train and test the network on a simulated dataset with only affine motion, achieving IoU and detection rate up to 71% and 87% respectively.
title EV-LayerSegNet: Self-supervised Motion Segmentation using Event Cameras
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.06596