Learning Temporally Consistent Video Depth from Video Diffusion Priors
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866916783175237632 |
|---|---|
| author | Shao, Jiahao Yang, Yuanbo Zhou, Hongyu Zhang, Youmin Shen, Yujun Guizilini, Vitor Wang, Yue Poggi, Matteo Liao, Yiyi |
| author_facet | Shao, Jiahao Yang, Yuanbo Zhou, Hongyu Zhang, Youmin Shen, Yujun Guizilini, Vitor Wang, Yue Poggi, Matteo Liao, Yiyi |
| contents | This work addresses the challenge of streamed video depth estimation, which expects not only per-frame accuracy but, more importantly, cross-frame consistency. We argue that sharing contextual information between frames or clips is pivotal in fostering temporal consistency. Therefore, we reformulate depth prediction into a conditional generation problem to provide contextual information within a clip and across clips. Specifically, we propose a consistent context-aware training and inference strategy for arbitrarily long videos to provide cross-clip context. We sample independent noise levels for each frame within a clip during training while using a sliding window strategy and initializing overlapping frames with previously predicted frames without adding noise. Moreover, we design an effective training strategy to provide context within a clip. Extensive experimental results validate our design choices and demonstrate the superiority of our approach, dubbed ChronoDepth. Project page: https://xdimlab.github.io/ChronoDepth/. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2406_01493 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Learning Temporally Consistent Video Depth from Video Diffusion Priors Shao, Jiahao Yang, Yuanbo Zhou, Hongyu Zhang, Youmin Shen, Yujun Guizilini, Vitor Wang, Yue Poggi, Matteo Liao, Yiyi Computer Vision and Pattern Recognition This work addresses the challenge of streamed video depth estimation, which expects not only per-frame accuracy but, more importantly, cross-frame consistency. We argue that sharing contextual information between frames or clips is pivotal in fostering temporal consistency. Therefore, we reformulate depth prediction into a conditional generation problem to provide contextual information within a clip and across clips. Specifically, we propose a consistent context-aware training and inference strategy for arbitrarily long videos to provide cross-clip context. We sample independent noise levels for each frame within a clip during training while using a sliding window strategy and initializing overlapping frames with previously predicted frames without adding noise. Moreover, we design an effective training strategy to provide context within a clip. Extensive experimental results validate our design choices and demonstrate the superiority of our approach, dubbed ChronoDepth. Project page: https://xdimlab.github.io/ChronoDepth/. |
| title | Learning Temporally Consistent Video Depth from Video Diffusion Priors |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2406.01493 |