The Dynamic Prior: Understanding 3D Structures for Casual Dynamic Videos

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Zhuoyuan, Yang, Xurui, Huang, Jiahui, Wang, Yue, Gao, Jun
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918233666224128
author Wu, Zhuoyuan
Yang, Xurui
Huang, Jiahui
Wang, Yue
Gao, Jun
author_facet Wu, Zhuoyuan
Yang, Xurui
Huang, Jiahui
Wang, Yue
Gao, Jun
contents Estimating accurate camera poses, 3D scene geometry, and object motion from in-the-wild videos is a long-standing challenge for classical structure from motion pipelines due to the presence of dynamic objects. Recent learning-based methods attempt to overcome this challenge by training motion estimators to filter dynamic objects and focus on the static background. However, their performance is largely limited by the availability of large-scale motion segmentation datasets, resulting in inaccurate segmentation and, therefore, inferior structural 3D understanding. In this work, we introduce the Dynamic Prior (\ourmodel) to robustly identify dynamic objects without task-specific training, leveraging the powerful reasoning capabilities of Vision-Language Models (VLMs) and the fine-grained spatial segmentation capacity of SAM2. \ourmodel can be seamlessly integrated into state-of-the-art pipelines for camera pose optimization, depth reconstruction, and 4D trajectory estimation. Extensive experiments on both synthetic and real-world videos demonstrate that \ourmodel not only achieves state-of-the-art performance on motion segmentation, but also significantly improves accuracy and robustness for structural 3D understanding.
format Preprint
id arxiv_https___arxiv_org_abs_2512_05398
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle The Dynamic Prior: Understanding 3D Structures for Casual Dynamic Videos
Wu, Zhuoyuan
Yang, Xurui
Huang, Jiahui
Wang, Yue
Gao, Jun
Computer Vision and Pattern Recognition
Estimating accurate camera poses, 3D scene geometry, and object motion from in-the-wild videos is a long-standing challenge for classical structure from motion pipelines due to the presence of dynamic objects. Recent learning-based methods attempt to overcome this challenge by training motion estimators to filter dynamic objects and focus on the static background. However, their performance is largely limited by the availability of large-scale motion segmentation datasets, resulting in inaccurate segmentation and, therefore, inferior structural 3D understanding. In this work, we introduce the Dynamic Prior (\ourmodel) to robustly identify dynamic objects without task-specific training, leveraging the powerful reasoning capabilities of Vision-Language Models (VLMs) and the fine-grained spatial segmentation capacity of SAM2. \ourmodel can be seamlessly integrated into state-of-the-art pipelines for camera pose optimization, depth reconstruction, and 4D trajectory estimation. Extensive experiments on both synthetic and real-world videos demonstrate that \ourmodel not only achieves state-of-the-art performance on motion segmentation, but also significantly improves accuracy and robustness for structural 3D understanding.
title The Dynamic Prior: Understanding 3D Structures for Casual Dynamic Videos
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.05398