MoSca: Dynamic Gaussian Fusion from Casual Videos via 4D Motion Scaffolds

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lei, Jiahui, Weng, Yijia, Harley, Adam, Guibas, Leonidas, Daniilidis, Kostas
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916499039453184
author Lei, Jiahui
Weng, Yijia
Harley, Adam
Guibas, Leonidas
Daniilidis, Kostas
author_facet Lei, Jiahui
Weng, Yijia
Harley, Adam
Guibas, Leonidas
Daniilidis, Kostas
contents We introduce 4D Motion Scaffolds (MoSca), a modern 4D reconstruction system designed to reconstruct and synthesize novel views of dynamic scenes from monocular videos captured casually in the wild. To address such a challenging and ill-posed inverse problem, we leverage prior knowledge from foundational vision models and lift the video data to a novel Motion Scaffold (MoSca) representation, which compactly and smoothly encodes the underlying motions/deformations. The scene geometry and appearance are then disentangled from the deformation field and are encoded by globally fusing the Gaussians anchored onto the MoSca and optimized via Gaussian Splatting. Additionally, camera focal length and poses can be solved using bundle adjustment without the need of any other pose estimation tools. Experiments demonstrate state-of-the-art performance on dynamic rendering benchmarks and its effectiveness on real videos.
format Preprint
id arxiv_https___arxiv_org_abs_2405_17421
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle MoSca: Dynamic Gaussian Fusion from Casual Videos via 4D Motion Scaffolds
Lei, Jiahui
Weng, Yijia
Harley, Adam
Guibas, Leonidas
Daniilidis, Kostas
Computer Vision and Pattern Recognition
Graphics
We introduce 4D Motion Scaffolds (MoSca), a modern 4D reconstruction system designed to reconstruct and synthesize novel views of dynamic scenes from monocular videos captured casually in the wild. To address such a challenging and ill-posed inverse problem, we leverage prior knowledge from foundational vision models and lift the video data to a novel Motion Scaffold (MoSca) representation, which compactly and smoothly encodes the underlying motions/deformations. The scene geometry and appearance are then disentangled from the deformation field and are encoded by globally fusing the Gaussians anchored onto the MoSca and optimized via Gaussian Splatting. Additionally, camera focal length and poses can be solved using bundle adjustment without the need of any other pose estimation tools. Experiments demonstrate state-of-the-art performance on dynamic rendering benchmarks and its effectiveness on real videos.
title MoSca: Dynamic Gaussian Fusion from Casual Videos via 4D Motion Scaffolds
topic Computer Vision and Pattern Recognition
Graphics
url https://arxiv.org/abs/2405.17421