VideoLifter: Lifting Videos to 3D with Fast Hierarchical Stereo Alignment

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cong, Wenyan, Zhu, Hanqing, Wang, Kevin, Lei, Jiahui, Stearns, Colton, Cai, Yuanhao, Guibas, Leonidas, Wang, Zhangyang, Fan, Zhiwen
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913930255794176
author Cong, Wenyan
Zhu, Hanqing
Wang, Kevin
Lei, Jiahui
Stearns, Colton
Cai, Yuanhao
Guibas, Leonidas
Wang, Zhangyang
Fan, Zhiwen
author_facet Cong, Wenyan
Zhu, Hanqing
Wang, Kevin
Lei, Jiahui
Stearns, Colton
Cai, Yuanhao
Guibas, Leonidas
Wang, Zhangyang
Fan, Zhiwen
contents Efficiently reconstructing 3D scenes from monocular video remains a core challenge in computer vision, vital for applications in virtual reality, robotics, and scene understanding. Recently, frame-by-frame progressive reconstruction without camera poses is commonly adopted, incurring high computational overhead and compounding errors when scaling to longer videos. To overcome these issues, we introduce VideoLifter, a novel video-to-3D pipeline that leverages a local-to-global strategy on a fragment basis, achieving both extreme efficiency and SOTA quality. Locally, VideoLifter leverages learnable 3D priors to register fragments, extracting essential information for subsequent 3D Gaussian initialization with enforced inter-fragment consistency and optimized efficiency. Globally, it employs a tree-based hierarchical merging method with key frame guidance for inter-fragment alignment, pairwise merging with Gaussian point pruning, and subsequent joint optimization to ensure global consistency while efficiently mitigating cumulative errors. This approach significantly accelerates the reconstruction process, reducing training time by over 82% while holding better visual quality than current SOTA methods.
format Preprint
id arxiv_https___arxiv_org_abs_2501_01949
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle VideoLifter: Lifting Videos to 3D with Fast Hierarchical Stereo Alignment
Cong, Wenyan
Zhu, Hanqing
Wang, Kevin
Lei, Jiahui
Stearns, Colton
Cai, Yuanhao
Guibas, Leonidas
Wang, Zhangyang
Fan, Zhiwen
Computer Vision and Pattern Recognition
Efficiently reconstructing 3D scenes from monocular video remains a core challenge in computer vision, vital for applications in virtual reality, robotics, and scene understanding. Recently, frame-by-frame progressive reconstruction without camera poses is commonly adopted, incurring high computational overhead and compounding errors when scaling to longer videos. To overcome these issues, we introduce VideoLifter, a novel video-to-3D pipeline that leverages a local-to-global strategy on a fragment basis, achieving both extreme efficiency and SOTA quality. Locally, VideoLifter leverages learnable 3D priors to register fragments, extracting essential information for subsequent 3D Gaussian initialization with enforced inter-fragment consistency and optimized efficiency. Globally, it employs a tree-based hierarchical merging method with key frame guidance for inter-fragment alignment, pairwise merging with Gaussian point pruning, and subsequent joint optimization to ensure global consistency while efficiently mitigating cumulative errors. This approach significantly accelerates the reconstruction process, reducing training time by over 82% while holding better visual quality than current SOTA methods.
title VideoLifter: Lifting Videos to 3D with Fast Hierarchical Stereo Alignment
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2501.01949