ShapeGaussian: High-Fidelity 4D Human Reconstruction in Monocular Videos via Vision Priors

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liang, Zhenxiao, Zhang, Ning, Tang, Youbao, Lin, Ruei-Sung, Huang, Qixing, Chang, Peng, Xiao, Jing
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915777777500160
author Liang, Zhenxiao
Zhang, Ning
Tang, Youbao
Lin, Ruei-Sung
Huang, Qixing
Chang, Peng
Xiao, Jing
author_facet Liang, Zhenxiao
Zhang, Ning
Tang, Youbao
Lin, Ruei-Sung
Huang, Qixing
Chang, Peng
Xiao, Jing
contents We introduce ShapeGaussian, a high-fidelity, template-free method for 4D human reconstruction from casual monocular videos. Generic reconstruction methods lacking robust vision priors, such as 4DGS, struggle to capture high-deformation human motion without multi-view cues. While template-based approaches, primarily relying on SMPL, such as HUGS, can produce photorealistic results, they are highly susceptible to errors in human pose estimation, often leading to unrealistic artifacts. In contrast, ShapeGaussian effectively integrates template-free vision priors to achieve both high-fidelity and robust scene reconstructions. Our method follows a two-step pipeline: first, we learn a coarse, deformable geometry using pretrained models that estimate data-driven priors, providing a foundation for reconstruction. Then, we refine this geometry using a neural deformation model to capture fine-grained dynamic details. By leveraging 2D vision priors, we mitigate artifacts from erroneous pose estimation in template-based methods and employ multiple reference frames to resolve the invisibility issue of 2D keypoints in a template-free manner. Extensive experiments demonstrate that ShapeGaussian surpasses template-based methods in reconstruction accuracy, achieving superior visual quality and robustness across diverse human motions in casual monocular videos.
format Preprint
id arxiv_https___arxiv_org_abs_2602_05572
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle ShapeGaussian: High-Fidelity 4D Human Reconstruction in Monocular Videos via Vision Priors
Liang, Zhenxiao
Zhang, Ning
Tang, Youbao
Lin, Ruei-Sung
Huang, Qixing
Chang, Peng
Xiao, Jing
Computer Vision and Pattern Recognition
We introduce ShapeGaussian, a high-fidelity, template-free method for 4D human reconstruction from casual monocular videos. Generic reconstruction methods lacking robust vision priors, such as 4DGS, struggle to capture high-deformation human motion without multi-view cues. While template-based approaches, primarily relying on SMPL, such as HUGS, can produce photorealistic results, they are highly susceptible to errors in human pose estimation, often leading to unrealistic artifacts. In contrast, ShapeGaussian effectively integrates template-free vision priors to achieve both high-fidelity and robust scene reconstructions. Our method follows a two-step pipeline: first, we learn a coarse, deformable geometry using pretrained models that estimate data-driven priors, providing a foundation for reconstruction. Then, we refine this geometry using a neural deformation model to capture fine-grained dynamic details. By leveraging 2D vision priors, we mitigate artifacts from erroneous pose estimation in template-based methods and employ multiple reference frames to resolve the invisibility issue of 2D keypoints in a template-free manner. Extensive experiments demonstrate that ShapeGaussian surpasses template-based methods in reconstruction accuracy, achieving superior visual quality and robustness across diverse human motions in casual monocular videos.
title ShapeGaussian: High-Fidelity 4D Human Reconstruction in Monocular Videos via Vision Priors
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2602.05572