Learning Monocular Depth from Events via Egomotion Compensation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Meng, Haitao, Zhong, Chonghao, Tang, Sheng, JunJia, Lian, Lin, Wenwei, Bing, Zhenshan, Chang, Yi, Chen, Gang, Knoll, Alois
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910764484263936
author Meng, Haitao
Zhong, Chonghao
Tang, Sheng
JunJia, Lian
Lin, Wenwei
Bing, Zhenshan
Chang, Yi
Chen, Gang
Knoll, Alois
author_facet Meng, Haitao
Zhong, Chonghao
Tang, Sheng
JunJia, Lian
Lin, Wenwei
Bing, Zhenshan
Chang, Yi
Chen, Gang
Knoll, Alois
contents Event cameras are neuromorphically inspired sensors that sparsely and asynchronously report brightness changes. Their unique characteristics of high temporal resolution, high dynamic range, and low power consumption make them well-suited for addressing challenges in monocular depth estimation (e.g., high-speed or low-lighting conditions). However, current existing methods primarily treat event streams as black-box learning systems without incorporating prior physical principles, thus becoming over-parameterized and failing to fully exploit the rich temporal information inherent in event camera data. To address this limitation, we incorporate physical motion principles to propose an interpretable monocular depth estimation framework, where the likelihood of various depth hypotheses is explicitly determined by the effect of motion compensation. To achieve this, we propose a Focus Cost Discrimination (FCD) module that measures the clarity of edges as an essential indicator of focus level and integrates spatial surroundings to facilitate cost estimation. Furthermore, we analyze the noise patterns within our framework and improve it with the newly introduced Inter-Hypotheses Cost Aggregation (IHCA) module, where the cost volume is refined through cost trend prediction and multi-scale cost consistency constraints. Extensive experiments on real-world and synthetic datasets demonstrate that our proposed framework outperforms cutting-edge methods by up to 10\% in terms of the absolute relative error metric, revealing superior performance in predicting accuracy.
format Preprint
id arxiv_https___arxiv_org_abs_2412_19067
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Learning Monocular Depth from Events via Egomotion Compensation
Meng, Haitao
Zhong, Chonghao
Tang, Sheng
JunJia, Lian
Lin, Wenwei
Bing, Zhenshan
Chang, Yi
Chen, Gang
Knoll, Alois
Computer Vision and Pattern Recognition
Machine Learning
Robotics
Event cameras are neuromorphically inspired sensors that sparsely and asynchronously report brightness changes. Their unique characteristics of high temporal resolution, high dynamic range, and low power consumption make them well-suited for addressing challenges in monocular depth estimation (e.g., high-speed or low-lighting conditions). However, current existing methods primarily treat event streams as black-box learning systems without incorporating prior physical principles, thus becoming over-parameterized and failing to fully exploit the rich temporal information inherent in event camera data. To address this limitation, we incorporate physical motion principles to propose an interpretable monocular depth estimation framework, where the likelihood of various depth hypotheses is explicitly determined by the effect of motion compensation. To achieve this, we propose a Focus Cost Discrimination (FCD) module that measures the clarity of edges as an essential indicator of focus level and integrates spatial surroundings to facilitate cost estimation. Furthermore, we analyze the noise patterns within our framework and improve it with the newly introduced Inter-Hypotheses Cost Aggregation (IHCA) module, where the cost volume is refined through cost trend prediction and multi-scale cost consistency constraints. Extensive experiments on real-world and synthetic datasets demonstrate that our proposed framework outperforms cutting-edge methods by up to 10\% in terms of the absolute relative error metric, revealing superior performance in predicting accuracy.
title Learning Monocular Depth from Events via Egomotion Compensation
topic Computer Vision and Pattern Recognition
Machine Learning
Robotics
url https://arxiv.org/abs/2412.19067