VideoArtGS: Building Digital Twins of Articulated Objects from Monocular Video

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Yu, Jia, Baoxiong, Lu, Ruijie, Gan, Chuyue, Chen, Huayu, Ni, Junfeng, Zhu, Song-Chun, Huang, Siyuan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914050606104576
author Liu, Yu
Jia, Baoxiong
Lu, Ruijie
Gan, Chuyue
Chen, Huayu
Ni, Junfeng
Zhu, Song-Chun
Huang, Siyuan
author_facet Liu, Yu
Jia, Baoxiong
Lu, Ruijie
Gan, Chuyue
Chen, Huayu
Ni, Junfeng
Zhu, Song-Chun
Huang, Siyuan
contents Building digital twins of articulated objects from monocular video presents an essential challenge in computer vision, which requires simultaneous reconstruction of object geometry, part segmentation, and articulation parameters from limited viewpoint inputs. Monocular video offers an attractive input format due to its simplicity and scalability; however, it's challenging to disentangle the object geometry and part dynamics with visual supervision alone, as the joint movement of the camera and parts leads to ill-posed estimation. While motion priors from pre-trained tracking models can alleviate the issue, how to effectively integrate them for articulation learning remains largely unexplored. To address this problem, we introduce VideoArtGS, a novel approach that reconstructs high-fidelity digital twins of articulated objects from monocular video. We propose a motion prior guidance pipeline that analyzes 3D tracks, filters noise, and provides reliable initialization of articulation parameters. We also design a hybrid center-grid part assignment module for articulation-based deformation fields that captures accurate part motion. VideoArtGS demonstrates state-of-the-art performance in articulation and mesh reconstruction, reducing the reconstruction error by about two orders of magnitude compared to existing methods. VideoArtGS enables practical digital twin creation from monocular video, establishing a new benchmark for video-based articulated object reconstruction. Our work is made publicly available at: https://videoartgs.github.io.
format Preprint
id arxiv_https___arxiv_org_abs_2509_17647
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle VideoArtGS: Building Digital Twins of Articulated Objects from Monocular Video
Liu, Yu
Jia, Baoxiong
Lu, Ruijie
Gan, Chuyue
Chen, Huayu
Ni, Junfeng
Zhu, Song-Chun
Huang, Siyuan
Computer Vision and Pattern Recognition
Artificial Intelligence
Robotics
Building digital twins of articulated objects from monocular video presents an essential challenge in computer vision, which requires simultaneous reconstruction of object geometry, part segmentation, and articulation parameters from limited viewpoint inputs. Monocular video offers an attractive input format due to its simplicity and scalability; however, it's challenging to disentangle the object geometry and part dynamics with visual supervision alone, as the joint movement of the camera and parts leads to ill-posed estimation. While motion priors from pre-trained tracking models can alleviate the issue, how to effectively integrate them for articulation learning remains largely unexplored. To address this problem, we introduce VideoArtGS, a novel approach that reconstructs high-fidelity digital twins of articulated objects from monocular video. We propose a motion prior guidance pipeline that analyzes 3D tracks, filters noise, and provides reliable initialization of articulation parameters. We also design a hybrid center-grid part assignment module for articulation-based deformation fields that captures accurate part motion. VideoArtGS demonstrates state-of-the-art performance in articulation and mesh reconstruction, reducing the reconstruction error by about two orders of magnitude compared to existing methods. VideoArtGS enables practical digital twin creation from monocular video, establishing a new benchmark for video-based articulated object reconstruction. Our work is made publicly available at: https://videoartgs.github.io.
title VideoArtGS: Building Digital Twins of Articulated Objects from Monocular Video
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Robotics
url https://arxiv.org/abs/2509.17647