Uni4D: Unifying Visual Foundation Models for 4D Modeling from a Single Video

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yao, David Yifan, Zhai, Albert J., Wang, Shenlong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912297641836544
author Yao, David Yifan
Zhai, Albert J.
Wang, Shenlong
author_facet Yao, David Yifan
Zhai, Albert J.
Wang, Shenlong
contents This paper presents a unified approach to understanding dynamic scenes from casual videos. Large pretrained vision foundation models, such as vision-language, video depth prediction, motion tracking, and segmentation models, offer promising capabilities. However, training a single model for comprehensive 4D understanding remains challenging. We introduce Uni4D, a multi-stage optimization framework that harnesses multiple pretrained models to advance dynamic 3D modeling, including static/dynamic reconstruction, camera pose estimation, and dense 3D motion tracking. Our results show state-of-the-art performance in dynamic 4D modeling with superior visual quality. Notably, Uni4D requires no retraining or fine-tuning, highlighting the effectiveness of repurposing visual foundation models for 4D understanding.
format Preprint
id arxiv_https___arxiv_org_abs_2503_21761
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Uni4D: Unifying Visual Foundation Models for 4D Modeling from a Single Video
Yao, David Yifan
Zhai, Albert J.
Wang, Shenlong
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
This paper presents a unified approach to understanding dynamic scenes from casual videos. Large pretrained vision foundation models, such as vision-language, video depth prediction, motion tracking, and segmentation models, offer promising capabilities. However, training a single model for comprehensive 4D understanding remains challenging. We introduce Uni4D, a multi-stage optimization framework that harnesses multiple pretrained models to advance dynamic 3D modeling, including static/dynamic reconstruction, camera pose estimation, and dense 3D motion tracking. Our results show state-of-the-art performance in dynamic 4D modeling with superior visual quality. Notably, Uni4D requires no retraining or fine-tuning, highlighting the effectiveness of repurposing visual foundation models for 4D understanding.
title Uni4D: Unifying Visual Foundation Models for 4D Modeling from a Single Video
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2503.21761