DoubleTake: Geometry Guided Depth Estimation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sayed, Mohamed, Aleotti, Filippo, Watson, Jamie, Qureshi, Zawar, Garcia-Hernando, Guillermo, Brostow, Gabriel, Vicente, Sara, Firman, Michael
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913431283564544
author Sayed, Mohamed
Aleotti, Filippo
Watson, Jamie
Qureshi, Zawar
Garcia-Hernando, Guillermo
Brostow, Gabriel
Vicente, Sara
Firman, Michael
author_facet Sayed, Mohamed
Aleotti, Filippo
Watson, Jamie
Qureshi, Zawar
Garcia-Hernando, Guillermo
Brostow, Gabriel
Vicente, Sara
Firman, Michael
contents Estimating depth from a sequence of posed RGB images is a fundamental computer vision task, with applications in augmented reality, path planning etc. Prior work typically makes use of previous frames in a multi view stereo framework, relying on matching textures in a local neighborhood. In contrast, our model leverages historical predictions by giving the latest 3D geometry data as an extra input to our network. This self-generated geometric hint can encode information from areas of the scene not covered by the keyframes and it is more regularized when compared to individual predicted depth maps for previous frames. We introduce a Hint MLP which combines cost volume features with a hint of the prior geometry, rendered as a depth map from the current camera location, together with a measure of the confidence in the prior geometry. We demonstrate that our method, which can run at interactive speeds, achieves state-of-the-art estimates of depth and 3D scene reconstruction in both offline and incremental evaluation scenarios.
format Preprint
id arxiv_https___arxiv_org_abs_2406_18387
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle DoubleTake: Geometry Guided Depth Estimation
Sayed, Mohamed
Aleotti, Filippo
Watson, Jamie
Qureshi, Zawar
Garcia-Hernando, Guillermo
Brostow, Gabriel
Vicente, Sara
Firman, Michael
Computer Vision and Pattern Recognition
Machine Learning
Estimating depth from a sequence of posed RGB images is a fundamental computer vision task, with applications in augmented reality, path planning etc. Prior work typically makes use of previous frames in a multi view stereo framework, relying on matching textures in a local neighborhood. In contrast, our model leverages historical predictions by giving the latest 3D geometry data as an extra input to our network. This self-generated geometric hint can encode information from areas of the scene not covered by the keyframes and it is more regularized when compared to individual predicted depth maps for previous frames. We introduce a Hint MLP which combines cost volume features with a hint of the prior geometry, rendered as a depth map from the current camera location, together with a measure of the confidence in the prior geometry. We demonstrate that our method, which can run at interactive speeds, achieves state-of-the-art estimates of depth and 3D scene reconstruction in both offline and incremental evaluation scenarios.
title DoubleTake: Geometry Guided Depth Estimation
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2406.18387