FutureDepth: Learning to Predict the Future Improves Video Depth Estimation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yasarla, Rajeev, Singh, Manish Kumar, Cai, Hong, Shi, Yunxiao, Jeong, Jisoo, Zhu, Yinhao, Han, Shizhong, Garrepalli, Risheek, Porikli, Fatih
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909457297965056
author Yasarla, Rajeev
Singh, Manish Kumar
Cai, Hong
Shi, Yunxiao
Jeong, Jisoo
Zhu, Yinhao
Han, Shizhong
Garrepalli, Risheek
Porikli, Fatih
author_facet Yasarla, Rajeev
Singh, Manish Kumar
Cai, Hong
Shi, Yunxiao
Jeong, Jisoo
Zhu, Yinhao
Han, Shizhong
Garrepalli, Risheek
Porikli, Fatih
contents In this paper, we propose a novel video depth estimation approach, FutureDepth, which enables the model to implicitly leverage multi-frame and motion cues to improve depth estimation by making it learn to predict the future at training. More specifically, we propose a future prediction network, F-Net, which takes the features of multiple consecutive frames and is trained to predict multi-frame features one time step ahead iteratively. In this way, F-Net learns the underlying motion and correspondence information, and we incorporate its features into the depth decoding process. Additionally, to enrich the learning of multiframe correspondence cues, we further leverage a reconstruction network, R-Net, which is trained via adaptively masked auto-encoding of multiframe feature volumes. At inference time, both F-Net and R-Net are used to produce queries to work with the depth decoder, as well as a final refinement network. Through extensive experiments on several benchmarks, i.e., NYUDv2, KITTI, DDAD, and Sintel, which cover indoor, driving, and open-domain scenarios, we show that FutureDepth significantly improves upon baseline models, outperforms existing video depth estimation methods, and sets new state-of-the-art (SOTA) accuracy. Furthermore, FutureDepth is more efficient than existing SOTA video depth estimation models and has similar latencies when comparing to monocular models
format Preprint
id arxiv_https___arxiv_org_abs_2403_12953
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle FutureDepth: Learning to Predict the Future Improves Video Depth Estimation
Yasarla, Rajeev
Singh, Manish Kumar
Cai, Hong
Shi, Yunxiao
Jeong, Jisoo
Zhu, Yinhao
Han, Shizhong
Garrepalli, Risheek
Porikli, Fatih
Computer Vision and Pattern Recognition
In this paper, we propose a novel video depth estimation approach, FutureDepth, which enables the model to implicitly leverage multi-frame and motion cues to improve depth estimation by making it learn to predict the future at training. More specifically, we propose a future prediction network, F-Net, which takes the features of multiple consecutive frames and is trained to predict multi-frame features one time step ahead iteratively. In this way, F-Net learns the underlying motion and correspondence information, and we incorporate its features into the depth decoding process. Additionally, to enrich the learning of multiframe correspondence cues, we further leverage a reconstruction network, R-Net, which is trained via adaptively masked auto-encoding of multiframe feature volumes. At inference time, both F-Net and R-Net are used to produce queries to work with the depth decoder, as well as a final refinement network. Through extensive experiments on several benchmarks, i.e., NYUDv2, KITTI, DDAD, and Sintel, which cover indoor, driving, and open-domain scenarios, we show that FutureDepth significantly improves upon baseline models, outperforms existing video depth estimation methods, and sets new state-of-the-art (SOTA) accuracy. Furthermore, FutureDepth is more efficient than existing SOTA video depth estimation models and has similar latencies when comparing to monocular models
title FutureDepth: Learning to Predict the Future Improves Video Depth Estimation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2403.12953