Buffer Anytime: Zero-Shot Video Depth and Normal from Image Priors

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kuang, Zhengfei, Zhang, Tianyuan, Zhang, Kai, Tan, Hao, Bi, Sai, Hu, Yiwei, Xu, Zexiang, Hasan, Milos, Wetzstein, Gordon, Luan, Fujun
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912133880479744
author Kuang, Zhengfei
Zhang, Tianyuan
Zhang, Kai
Tan, Hao
Bi, Sai
Hu, Yiwei
Xu, Zexiang
Hasan, Milos
Wetzstein, Gordon
Luan, Fujun
author_facet Kuang, Zhengfei
Zhang, Tianyuan
Zhang, Kai
Tan, Hao
Bi, Sai
Hu, Yiwei
Xu, Zexiang
Hasan, Milos
Wetzstein, Gordon
Luan, Fujun
contents We present Buffer Anytime, a framework for estimation of depth and normal maps (which we call geometric buffers) from video that eliminates the need for paired video--depth and video--normal training data. Instead of relying on large-scale annotated video datasets, we demonstrate high-quality video buffer estimation by leveraging single-image priors with temporal consistency constraints. Our zero-shot training strategy combines state-of-the-art image estimation models based on optical flow smoothness through a hybrid loss function, implemented via a lightweight temporal attention architecture. Applied to leading image models like Depth Anything V2 and Marigold-E2E-FT, our approach significantly improves temporal consistency while maintaining accuracy. Experiments show that our method not only outperforms image-based approaches but also achieves results comparable to state-of-the-art video models trained on large-scale paired video datasets, despite using no such paired video data.
format Preprint
id arxiv_https___arxiv_org_abs_2411_17249
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Buffer Anytime: Zero-Shot Video Depth and Normal from Image Priors
Kuang, Zhengfei
Zhang, Tianyuan
Zhang, Kai
Tan, Hao
Bi, Sai
Hu, Yiwei
Xu, Zexiang
Hasan, Milos
Wetzstein, Gordon
Luan, Fujun
Computer Vision and Pattern Recognition
Artificial Intelligence
We present Buffer Anytime, a framework for estimation of depth and normal maps (which we call geometric buffers) from video that eliminates the need for paired video--depth and video--normal training data. Instead of relying on large-scale annotated video datasets, we demonstrate high-quality video buffer estimation by leveraging single-image priors with temporal consistency constraints. Our zero-shot training strategy combines state-of-the-art image estimation models based on optical flow smoothness through a hybrid loss function, implemented via a lightweight temporal attention architecture. Applied to leading image models like Depth Anything V2 and Marigold-E2E-FT, our approach significantly improves temporal consistency while maintaining accuracy. Experiments show that our method not only outperforms image-based approaches but also achieves results comparable to state-of-the-art video models trained on large-scale paired video datasets, despite using no such paired video data.
title Buffer Anytime: Zero-Shot Video Depth and Normal from Image Priors
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2411.17249