Learning Temporally Consistent Video Depth from Video Diffusion Priors

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Shao, Jiahao, Yang, Yuanbo, Zhou, Hongyu, Zhang, Youmin, Shen, Yujun, Guizilini, Vitor, Wang, Yue, Poggi, Matteo, Liao, Yiyi
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916783175237632
author Shao, Jiahao
Yang, Yuanbo
Zhou, Hongyu
Zhang, Youmin
Shen, Yujun
Guizilini, Vitor
Wang, Yue
Poggi, Matteo
Liao, Yiyi
author_facet Shao, Jiahao
Yang, Yuanbo
Zhou, Hongyu
Zhang, Youmin
Shen, Yujun
Guizilini, Vitor
Wang, Yue
Poggi, Matteo
Liao, Yiyi
contents This work addresses the challenge of streamed video depth estimation, which expects not only per-frame accuracy but, more importantly, cross-frame consistency. We argue that sharing contextual information between frames or clips is pivotal in fostering temporal consistency. Therefore, we reformulate depth prediction into a conditional generation problem to provide contextual information within a clip and across clips. Specifically, we propose a consistent context-aware training and inference strategy for arbitrarily long videos to provide cross-clip context. We sample independent noise levels for each frame within a clip during training while using a sliding window strategy and initializing overlapping frames with previously predicted frames without adding noise. Moreover, we design an effective training strategy to provide context within a clip. Extensive experimental results validate our design choices and demonstrate the superiority of our approach, dubbed ChronoDepth. Project page: https://xdimlab.github.io/ChronoDepth/.
format Preprint
id arxiv_https___arxiv_org_abs_2406_01493
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Learning Temporally Consistent Video Depth from Video Diffusion Priors
Shao, Jiahao
Yang, Yuanbo
Zhou, Hongyu
Zhang, Youmin
Shen, Yujun
Guizilini, Vitor
Wang, Yue
Poggi, Matteo
Liao, Yiyi
Computer Vision and Pattern Recognition
This work addresses the challenge of streamed video depth estimation, which expects not only per-frame accuracy but, more importantly, cross-frame consistency. We argue that sharing contextual information between frames or clips is pivotal in fostering temporal consistency. Therefore, we reformulate depth prediction into a conditional generation problem to provide contextual information within a clip and across clips. Specifically, we propose a consistent context-aware training and inference strategy for arbitrarily long videos to provide cross-clip context. We sample independent noise levels for each frame within a clip during training while using a sliding window strategy and initializing overlapping frames with previously predicted frames without adding noise. Moreover, we design an effective training strategy to provide context within a clip. Extensive experimental results validate our design choices and demonstrate the superiority of our approach, dubbed ChronoDepth. Project page: https://xdimlab.github.io/ChronoDepth/.
title Learning Temporally Consistent Video Depth from Video Diffusion Priors
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2406.01493