4DVGGT-D: 4D Visual Geometry Transformer with Improved Dynamic Depth Estimation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zang, Ying, Liu, Xuanyi, Han, Yidong, Ji, Deyi, Ding, Chaotao, Hu, Yuanqi, Zhu, Qi, Li, Xuanfu, Ma, Jin, Sun, Lingyun, Chen, Tianrun, Zhu, Lanyun
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910212167827456
author Zang, Ying
Liu, Xuanyi
Han, Yidong
Ji, Deyi
Ding, Chaotao
Hu, Yuanqi
Zhu, Qi
Li, Xuanfu
Ma, Jin
Sun, Lingyun
Chen, Tianrun
Zhu, Lanyun
author_facet Zang, Ying
Liu, Xuanyi
Han, Yidong
Ji, Deyi
Ding, Chaotao
Hu, Yuanqi
Zhu, Qi
Li, Xuanfu
Ma, Jin
Sun, Lingyun
Chen, Tianrun
Zhu, Lanyun
contents Reconstructing dynamic 4D scenes from monocular videos is a fundamental yet challenging task. While recent 3D foundation models provide strong geometric priors, their performance significantly degrades in dynamic environments. This degradation stems from a fundamental tension: the inherent coupling of camera ego-motion and object motion within global attention mechanisms. In this paper, we propose a novel, training-free progressive decoupling framework that disentangles dynamics from statics in a principled, coarse-to-fine manner. Our core insight is to resolve the tension by first stabilizing the camera pose, followed by geometric refinement. Specifically, our approach consists of three synergistic components: (1) a Dynamic-Mask-Guided Pose Decoupling module that isolates pose estimation from dynamic interference, yielding a stable motion-free reference frame; (2) a Topological Subspace Surgery mechanism that orthogonally decomposes the depth manifold, safely preserving dynamic objects while injecting refined, mask-aware geometry into static regions; and (3) an Information-Theoretic Confidence-Aware Fusion strategy that formulates depth integration as a heteroscedastic Bayesian inference problem, adaptively blending multi-pass predictions via inverse-variance weighting. Extensive experiments on standard 4D reconstruction benchmarks demonstrate that our method achieves consistent and substantial improvements across principal point-cloud metrics. Notably, our approach shows competitive performance in robust 4D scene reconstruction without requiring fine-tuning, suggesting the potential of mathematically grounded dynamic-static disentanglement.
format Preprint
id arxiv_https___arxiv_org_abs_2605_12027
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle 4DVGGT-D: 4D Visual Geometry Transformer with Improved Dynamic Depth Estimation
Zang, Ying
Liu, Xuanyi
Han, Yidong
Ji, Deyi
Ding, Chaotao
Hu, Yuanqi
Zhu, Qi
Li, Xuanfu
Ma, Jin
Sun, Lingyun
Chen, Tianrun
Zhu, Lanyun
Computer Vision and Pattern Recognition
Reconstructing dynamic 4D scenes from monocular videos is a fundamental yet challenging task. While recent 3D foundation models provide strong geometric priors, their performance significantly degrades in dynamic environments. This degradation stems from a fundamental tension: the inherent coupling of camera ego-motion and object motion within global attention mechanisms. In this paper, we propose a novel, training-free progressive decoupling framework that disentangles dynamics from statics in a principled, coarse-to-fine manner. Our core insight is to resolve the tension by first stabilizing the camera pose, followed by geometric refinement. Specifically, our approach consists of three synergistic components: (1) a Dynamic-Mask-Guided Pose Decoupling module that isolates pose estimation from dynamic interference, yielding a stable motion-free reference frame; (2) a Topological Subspace Surgery mechanism that orthogonally decomposes the depth manifold, safely preserving dynamic objects while injecting refined, mask-aware geometry into static regions; and (3) an Information-Theoretic Confidence-Aware Fusion strategy that formulates depth integration as a heteroscedastic Bayesian inference problem, adaptively blending multi-pass predictions via inverse-variance weighting. Extensive experiments on standard 4D reconstruction benchmarks demonstrate that our method achieves consistent and substantial improvements across principal point-cloud metrics. Notably, our approach shows competitive performance in robust 4D scene reconstruction without requiring fine-tuning, suggesting the potential of mathematically grounded dynamic-static disentanglement.
title 4DVGGT-D: 4D Visual Geometry Transformer with Improved Dynamic Depth Estimation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.12027