DGS-LRM: Real-Time Deformable 3D Gaussian Reconstruction From Monocular Videos

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lin, Chieh Hubert, Lv, Zhaoyang, Wu, Songyin, Xu, Zhen, Nguyen-Phuoc, Thu, Tseng, Hung-Yu, Straub, Julian, Khan, Numair, Xiao, Lei, Yang, Ming-Hsuan, Ren, Yuheng, Newcombe, Richard, Dong, Zhao, Li, Zhengqin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916790784753664
author Lin, Chieh Hubert
Lv, Zhaoyang
Wu, Songyin
Xu, Zhen
Nguyen-Phuoc, Thu
Tseng, Hung-Yu
Straub, Julian
Khan, Numair
Xiao, Lei
Yang, Ming-Hsuan
Ren, Yuheng
Newcombe, Richard
Dong, Zhao
Li, Zhengqin
author_facet Lin, Chieh Hubert
Lv, Zhaoyang
Wu, Songyin
Xu, Zhen
Nguyen-Phuoc, Thu
Tseng, Hung-Yu
Straub, Julian
Khan, Numair
Xiao, Lei
Yang, Ming-Hsuan
Ren, Yuheng
Newcombe, Richard
Dong, Zhao
Li, Zhengqin
contents We introduce the Deformable Gaussian Splats Large Reconstruction Model (DGS-LRM), the first feed-forward method predicting deformable 3D Gaussian splats from a monocular posed video of any dynamic scene. Feed-forward scene reconstruction has gained significant attention for its ability to rapidly create digital replicas of real-world environments. However, most existing models are limited to static scenes and fail to reconstruct the motion of moving objects. Developing a feed-forward model for dynamic scene reconstruction poses significant challenges, including the scarcity of training data and the need for appropriate 3D representations and training paradigms. To address these challenges, we introduce several key technical contributions: an enhanced large-scale synthetic dataset with ground-truth multi-view videos and dense 3D scene flow supervision; a per-pixel deformable 3D Gaussian representation that is easy to learn, supports high-quality dynamic view synthesis, and enables long-range 3D tracking; and a large transformer network that achieves real-time, generalizable dynamic scene reconstruction. Extensive qualitative and quantitative experiments demonstrate that DGS-LRM achieves dynamic scene reconstruction quality comparable to optimization-based methods, while significantly outperforming the state-of-the-art predictive dynamic reconstruction method on real-world examples. Its predicted physically grounded 3D deformation is accurate and can readily adapt for long-range 3D tracking tasks, achieving performance on par with state-of-the-art monocular video 3D tracking methods.
format Preprint
id arxiv_https___arxiv_org_abs_2506_09997
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DGS-LRM: Real-Time Deformable 3D Gaussian Reconstruction From Monocular Videos
Lin, Chieh Hubert
Lv, Zhaoyang
Wu, Songyin
Xu, Zhen
Nguyen-Phuoc, Thu
Tseng, Hung-Yu
Straub, Julian
Khan, Numair
Xiao, Lei
Yang, Ming-Hsuan
Ren, Yuheng
Newcombe, Richard
Dong, Zhao
Li, Zhengqin
Graphics
Artificial Intelligence
Computer Vision and Pattern Recognition
Machine Learning
We introduce the Deformable Gaussian Splats Large Reconstruction Model (DGS-LRM), the first feed-forward method predicting deformable 3D Gaussian splats from a monocular posed video of any dynamic scene. Feed-forward scene reconstruction has gained significant attention for its ability to rapidly create digital replicas of real-world environments. However, most existing models are limited to static scenes and fail to reconstruct the motion of moving objects. Developing a feed-forward model for dynamic scene reconstruction poses significant challenges, including the scarcity of training data and the need for appropriate 3D representations and training paradigms. To address these challenges, we introduce several key technical contributions: an enhanced large-scale synthetic dataset with ground-truth multi-view videos and dense 3D scene flow supervision; a per-pixel deformable 3D Gaussian representation that is easy to learn, supports high-quality dynamic view synthesis, and enables long-range 3D tracking; and a large transformer network that achieves real-time, generalizable dynamic scene reconstruction. Extensive qualitative and quantitative experiments demonstrate that DGS-LRM achieves dynamic scene reconstruction quality comparable to optimization-based methods, while significantly outperforming the state-of-the-art predictive dynamic reconstruction method on real-world examples. Its predicted physically grounded 3D deformation is accurate and can readily adapt for long-range 3D tracking tasks, achieving performance on par with state-of-the-art monocular video 3D tracking methods.
title DGS-LRM: Real-Time Deformable 3D Gaussian Reconstruction From Monocular Videos
topic Graphics
Artificial Intelligence
Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2506.09997