Neuromorphic Monocular Depth Estimation with Uncertainty Modeling

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bergkvist, Viktor, Rydell, Felix, Forssén, Per-Erik, Gustafsson, David, Rideg, Johan
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918494638964736
author Bergkvist, Viktor
Rydell, Felix
Forssén, Per-Erik
Gustafsson, David
Rideg, Johan
author_facet Bergkvist, Viktor
Rydell, Felix
Forssén, Per-Erik
Gustafsson, David
Rideg, Johan
contents Event cameras offer distinct advantages over conventional frame-based sensors, including microsecond-level temporal resolution, high dynamic range, and low bandwidth. In this paper, we predict per-pixel depth distributions from monocular event streams using deep neural networks. We estimate uncertainty using Gaussian, log-normal, and evidential learning frameworks. We compare six event representations: spatio-temporal voxel grids with 1, 5, 10, and 20 temporal bins, the Compact Spatio-Temporal Representation (CSTR), and Time-Ordered Recent Event (TORE) volumes. Our U-Net-based models are trained on synthetic data and then fine-tuned on real sequences. We evaluate performance using absolute relative error, root mean squared error, and the area under the sparsification error. Quantitative results show that the representations perform similarly, while 10 bin log-normal and 5 bin evidential learning perform best across metrics. Our experiments demonstrate that uncertainty estimation can be successfully integrated into event-based monocular depth estimation, and be used to indicate pixels with reliable depth.
format Preprint
id arxiv_https___arxiv_org_abs_2605_10675
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Neuromorphic Monocular Depth Estimation with Uncertainty Modeling
Bergkvist, Viktor
Rydell, Felix
Forssén, Per-Erik
Gustafsson, David
Rideg, Johan
Computer Vision and Pattern Recognition
I.4.8; I.2.10; I.2.6
Event cameras offer distinct advantages over conventional frame-based sensors, including microsecond-level temporal resolution, high dynamic range, and low bandwidth. In this paper, we predict per-pixel depth distributions from monocular event streams using deep neural networks. We estimate uncertainty using Gaussian, log-normal, and evidential learning frameworks. We compare six event representations: spatio-temporal voxel grids with 1, 5, 10, and 20 temporal bins, the Compact Spatio-Temporal Representation (CSTR), and Time-Ordered Recent Event (TORE) volumes. Our U-Net-based models are trained on synthetic data and then fine-tuned on real sequences. We evaluate performance using absolute relative error, root mean squared error, and the area under the sparsification error. Quantitative results show that the representations perform similarly, while 10 bin log-normal and 5 bin evidential learning perform best across metrics. Our experiments demonstrate that uncertainty estimation can be successfully integrated into event-based monocular depth estimation, and be used to indicate pixels with reliable depth.
title Neuromorphic Monocular Depth Estimation with Uncertainty Modeling
topic Computer Vision and Pattern Recognition
I.4.8; I.2.10; I.2.6
url https://arxiv.org/abs/2605.10675