Buffer Anytime: Zero-Shot Video Depth and Normal from Image Priors
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912133880479744 |
|---|---|
| author | Kuang, Zhengfei Zhang, Tianyuan Zhang, Kai Tan, Hao Bi, Sai Hu, Yiwei Xu, Zexiang Hasan, Milos Wetzstein, Gordon Luan, Fujun |
| author_facet | Kuang, Zhengfei Zhang, Tianyuan Zhang, Kai Tan, Hao Bi, Sai Hu, Yiwei Xu, Zexiang Hasan, Milos Wetzstein, Gordon Luan, Fujun |
| contents | We present Buffer Anytime, a framework for estimation of depth and normal maps (which we call geometric buffers) from video that eliminates the need for paired video--depth and video--normal training data. Instead of relying on large-scale annotated video datasets, we demonstrate high-quality video buffer estimation by leveraging single-image priors with temporal consistency constraints. Our zero-shot training strategy combines state-of-the-art image estimation models based on optical flow smoothness through a hybrid loss function, implemented via a lightweight temporal attention architecture. Applied to leading image models like Depth Anything V2 and Marigold-E2E-FT, our approach significantly improves temporal consistency while maintaining accuracy. Experiments show that our method not only outperforms image-based approaches but also achieves results comparable to state-of-the-art video models trained on large-scale paired video datasets, despite using no such paired video data. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2411_17249 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Buffer Anytime: Zero-Shot Video Depth and Normal from Image Priors Kuang, Zhengfei Zhang, Tianyuan Zhang, Kai Tan, Hao Bi, Sai Hu, Yiwei Xu, Zexiang Hasan, Milos Wetzstein, Gordon Luan, Fujun Computer Vision and Pattern Recognition Artificial Intelligence We present Buffer Anytime, a framework for estimation of depth and normal maps (which we call geometric buffers) from video that eliminates the need for paired video--depth and video--normal training data. Instead of relying on large-scale annotated video datasets, we demonstrate high-quality video buffer estimation by leveraging single-image priors with temporal consistency constraints. Our zero-shot training strategy combines state-of-the-art image estimation models based on optical flow smoothness through a hybrid loss function, implemented via a lightweight temporal attention architecture. Applied to leading image models like Depth Anything V2 and Marigold-E2E-FT, our approach significantly improves temporal consistency while maintaining accuracy. Experiments show that our method not only outperforms image-based approaches but also achieves results comparable to state-of-the-art video models trained on large-scale paired video datasets, despite using no such paired video data. |
| title | Buffer Anytime: Zero-Shot Video Depth and Normal from Image Priors |
| topic | Computer Vision and Pattern Recognition Artificial Intelligence |
| url | https://arxiv.org/abs/2411.17249 |