RiGS: Rigid-aware 4D Gaussian Splatting from a Single Monocular Video

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Chenyu, Li, Wanhua, Chen, Zhu-Tian, Pfister, Hanspeter
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918518619897856
author Wu, Chenyu
Li, Wanhua
Chen, Zhu-Tian
Pfister, Hanspeter
author_facet Wu, Chenyu
Li, Wanhua
Chen, Zhu-Tian
Pfister, Hanspeter
contents Reconstructing dynamic 3D scenes from monocular videos is a fundamental yet highly challenging task, as real-world motions often involve both long-term smooth transformations and short-term complex deformations. Existing methods either struggle to maintain temporal consistency or fail to capture high-frequency dynamics due to limited motion modeling capacity. In this work, we present Rigid-aware 4D Gaussian Splatting (RiGS), which simultaneously captures motions across multiple temporal scales. Specifically, RiGS introduces three types of Gaussian primitives: static, rigid, and transient, which represent static backgrounds, long-term low-frequency motions, and short-term high-frequency dynamics, respectively. An object-wise dynamic mask is proposed to aggregate long-range spatiotemporal motion information and guide the decomposition of static and dynamic regions. To jointly model motion across scales, rigid Gaussians are allowed to transition into transient Gaussians based on their temporal duration, and both are optimized under scene flow guidance, providing dense 3D motion supervision. Extensive experiments demonstrate that RiGS achieves state-of-the-art performance on novel view synthesis benchmarks. Code is available at \hyperlink{https://github.com/ladvu/RiGS}{https://github.com/ladvu/RiGS}.
format Preprint
id arxiv_https___arxiv_org_abs_2605_23672
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle RiGS: Rigid-aware 4D Gaussian Splatting from a Single Monocular Video
Wu, Chenyu
Li, Wanhua
Chen, Zhu-Tian
Pfister, Hanspeter
Computer Vision and Pattern Recognition
Reconstructing dynamic 3D scenes from monocular videos is a fundamental yet highly challenging task, as real-world motions often involve both long-term smooth transformations and short-term complex deformations. Existing methods either struggle to maintain temporal consistency or fail to capture high-frequency dynamics due to limited motion modeling capacity. In this work, we present Rigid-aware 4D Gaussian Splatting (RiGS), which simultaneously captures motions across multiple temporal scales. Specifically, RiGS introduces three types of Gaussian primitives: static, rigid, and transient, which represent static backgrounds, long-term low-frequency motions, and short-term high-frequency dynamics, respectively. An object-wise dynamic mask is proposed to aggregate long-range spatiotemporal motion information and guide the decomposition of static and dynamic regions. To jointly model motion across scales, rigid Gaussians are allowed to transition into transient Gaussians based on their temporal duration, and both are optimized under scene flow guidance, providing dense 3D motion supervision. Extensive experiments demonstrate that RiGS achieves state-of-the-art performance on novel view synthesis benchmarks. Code is available at \hyperlink{https://github.com/ladvu/RiGS}{https://github.com/ladvu/RiGS}.
title RiGS: Rigid-aware 4D Gaussian Splatting from a Single Monocular Video
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.23672