Robust 4D Visual Geometry Transformer with Uncertainty-Aware Priors

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zang, Ying, Han, Yidong, Ding, Chaotao, Hu, Yuanqi, Ji, Deyi, Zhu, Qi, Li, Xuanfu, Ma, Jin, Sun, Lingyun, Chen, Tianrun, Zhu, Lanyun
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915930238353408
author Zang, Ying
Han, Yidong
Ding, Chaotao
Hu, Yuanqi
Ji, Deyi
Zhu, Qi
Li, Xuanfu
Ma, Jin
Sun, Lingyun
Chen, Tianrun
Zhu, Lanyun
author_facet Zang, Ying
Han, Yidong
Ding, Chaotao
Hu, Yuanqi
Ji, Deyi
Zhu, Qi
Li, Xuanfu
Ma, Jin
Sun, Lingyun
Chen, Tianrun
Zhu, Lanyun
contents Reconstructing dynamic 4D scenes is an important yet challenging task. While 3D foundation models like VGGT excel in static settings, they often struggle with dynamic sequences where motion causes significant geometric ambiguity. To address this, we present a framework designed to disentangle dynamic and static components by modeling uncertainty across different stages of the reconstruction process. Our approach introduces three synergistic mechanisms: (1) Entropy-Guided Subspace Projection, which leverages information-theoretic weighting to adaptively aggregate multi-head attention distributions, effectively isolating dynamic motion cues from semantic noise; (2) Local-Consistency Driven Geometry Purification, which enforces spatial continuity via radius-based neighborhood constraints to eliminate structural outliers; and (3) Uncertainty-Aware Cross-View Consistency, which formulates multi-view projection refinement as a heteroscedastic maximum likelihood estimation problem, utilizing depth confidence as a probabilistic weight. Experiments on dynamic benchmarks show that our approach outperforms current state-of-the-art methods, reducing Mean Accuracy error by 13.43\% and improving segmentation F-measure by 10.49\%. Our framework maintains the efficiency of feed-forward inference and requires no task-specific fine-tuning or per-scene optimization.
format Preprint
id arxiv_https___arxiv_org_abs_2604_09366
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Robust 4D Visual Geometry Transformer with Uncertainty-Aware Priors
Zang, Ying
Han, Yidong
Ding, Chaotao
Hu, Yuanqi
Ji, Deyi
Zhu, Qi
Li, Xuanfu
Ma, Jin
Sun, Lingyun
Chen, Tianrun
Zhu, Lanyun
Computer Vision and Pattern Recognition
Reconstructing dynamic 4D scenes is an important yet challenging task. While 3D foundation models like VGGT excel in static settings, they often struggle with dynamic sequences where motion causes significant geometric ambiguity. To address this, we present a framework designed to disentangle dynamic and static components by modeling uncertainty across different stages of the reconstruction process. Our approach introduces three synergistic mechanisms: (1) Entropy-Guided Subspace Projection, which leverages information-theoretic weighting to adaptively aggregate multi-head attention distributions, effectively isolating dynamic motion cues from semantic noise; (2) Local-Consistency Driven Geometry Purification, which enforces spatial continuity via radius-based neighborhood constraints to eliminate structural outliers; and (3) Uncertainty-Aware Cross-View Consistency, which formulates multi-view projection refinement as a heteroscedastic maximum likelihood estimation problem, utilizing depth confidence as a probabilistic weight. Experiments on dynamic benchmarks show that our approach outperforms current state-of-the-art methods, reducing Mean Accuracy error by 13.43\% and improving segmentation F-measure by 10.49\%. Our framework maintains the efficiency of feed-forward inference and requires no task-specific fine-tuning or per-scene optimization.
title Robust 4D Visual Geometry Transformer with Uncertainty-Aware Priors
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2604.09366