Learning Monocular Depth from Focus with Event Focal Stack

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jiang, Chenxu, Lin, Mingyuan, Zhang, Chi, Wang, Zhenghai, Yu, Lei
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929340606840832
author Jiang, Chenxu
Lin, Mingyuan
Zhang, Chi
Wang, Zhenghai
Yu, Lei
author_facet Jiang, Chenxu
Lin, Mingyuan
Zhang, Chi
Wang, Zhenghai
Yu, Lei
contents Depth from Focus estimates depth by determining the moment of maximum focus from multiple shots at different focal distances, i.e. the Focal Stack. However, the limited sampling rate of conventional optical cameras makes it difficult to obtain sufficient focus cues during the focal sweep. Inspired by biological vision, the event camera records intensity changes over time in extremely low latency, which provides more temporal information for focus time acquisition. In this study, we propose the EDFF Network to estimate sparse depth from the Event Focal Stack. Specifically, we utilize the event voxel grid to encode intensity change information and project event time surface into the depth domain to preserve per-pixel focal distance information. A Focal-Distance-guided Cross-Modal Attention Module is presented to fuse the information mentioned above. Additionally, we propose a Multi-level Depth Fusion Block designed to integrate results from each level of a UNet-like architecture and produce the final output. Extensive experiments validate that our method outperforms existing state-of-the-art approaches.
format Preprint
id arxiv_https___arxiv_org_abs_2405_06944
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Learning Monocular Depth from Focus with Event Focal Stack
Jiang, Chenxu
Lin, Mingyuan
Zhang, Chi
Wang, Zhenghai
Yu, Lei
Computer Vision and Pattern Recognition
Depth from Focus estimates depth by determining the moment of maximum focus from multiple shots at different focal distances, i.e. the Focal Stack. However, the limited sampling rate of conventional optical cameras makes it difficult to obtain sufficient focus cues during the focal sweep. Inspired by biological vision, the event camera records intensity changes over time in extremely low latency, which provides more temporal information for focus time acquisition. In this study, we propose the EDFF Network to estimate sparse depth from the Event Focal Stack. Specifically, we utilize the event voxel grid to encode intensity change information and project event time surface into the depth domain to preserve per-pixel focal distance information. A Focal-Distance-guided Cross-Modal Attention Module is presented to fuse the information mentioned above. Additionally, we propose a Multi-level Depth Fusion Block designed to integrate results from each level of a UNet-like architecture and produce the final output. Extensive experiments validate that our method outperforms existing state-of-the-art approaches.
title Learning Monocular Depth from Focus with Event Focal Stack
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2405.06944