UniSH: Unifying Scene and Human Reconstruction in a Feed-Forward Pass

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Mengfei, Li, Peng, Zhang, Zheng, Lu, Jiahao, Zhao, Chengfeng, Xue, Wei, Liu, Qifeng, Peng, Sida, Zhang, Wenxiao, Luo, Wenhan, Liu, Yuan, Guo, Yike
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917182967906304
author Li, Mengfei
Li, Peng
Zhang, Zheng
Lu, Jiahao
Zhao, Chengfeng
Xue, Wei
Liu, Qifeng
Peng, Sida
Zhang, Wenxiao
Luo, Wenhan
Liu, Yuan
Guo, Yike
author_facet Li, Mengfei
Li, Peng
Zhang, Zheng
Lu, Jiahao
Zhao, Chengfeng
Xue, Wei
Liu, Qifeng
Peng, Sida
Zhang, Wenxiao
Luo, Wenhan
Liu, Yuan
Guo, Yike
contents We present UniSH, a unified, feed-forward framework for joint metric-scale 3D scene and human reconstruction. A key challenge in this domain is the scarcity of large-scale, annotated real-world data, forcing a reliance on synthetic datasets. This reliance introduces a significant sim-to-real domain gap, leading to poor generalization, low-fidelity human geometry, and poor alignment on in-the-wild videos. To address this, we propose an innovative training paradigm that effectively leverages unlabeled in-the-wild data. Our framework bridges strong, disparate priors from scene reconstruction and HMR, and is trained with two core components: (1) a robust distillation strategy to refine human surface details by distilling high-frequency details from an expert depth model, and (2) a two-stage supervision scheme, which first learns coarse localization on synthetic data, then fine-tunes on real data by directly optimizing the geometric correspondence between the SMPL mesh and the human point cloud. This approach enables our feed-forward model to jointly recover high-fidelity scene geometry, human point clouds, camera parameters, and coherent, metric-scale SMPL bodies, all in a single forward pass. Extensive experiments demonstrate that our model achieves state-of-the-art performance on human-centric scene reconstruction and delivers highly competitive results on global human motion estimation, comparing favorably against both optimization-based frameworks and HMR-only methods. Project page: https://murphylmf.github.io/UniSH/
format Preprint
id arxiv_https___arxiv_org_abs_2601_01222
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle UniSH: Unifying Scene and Human Reconstruction in a Feed-Forward Pass
Li, Mengfei
Li, Peng
Zhang, Zheng
Lu, Jiahao
Zhao, Chengfeng
Xue, Wei
Liu, Qifeng
Peng, Sida
Zhang, Wenxiao
Luo, Wenhan
Liu, Yuan
Guo, Yike
Computer Vision and Pattern Recognition
We present UniSH, a unified, feed-forward framework for joint metric-scale 3D scene and human reconstruction. A key challenge in this domain is the scarcity of large-scale, annotated real-world data, forcing a reliance on synthetic datasets. This reliance introduces a significant sim-to-real domain gap, leading to poor generalization, low-fidelity human geometry, and poor alignment on in-the-wild videos. To address this, we propose an innovative training paradigm that effectively leverages unlabeled in-the-wild data. Our framework bridges strong, disparate priors from scene reconstruction and HMR, and is trained with two core components: (1) a robust distillation strategy to refine human surface details by distilling high-frequency details from an expert depth model, and (2) a two-stage supervision scheme, which first learns coarse localization on synthetic data, then fine-tunes on real data by directly optimizing the geometric correspondence between the SMPL mesh and the human point cloud. This approach enables our feed-forward model to jointly recover high-fidelity scene geometry, human point clouds, camera parameters, and coherent, metric-scale SMPL bodies, all in a single forward pass. Extensive experiments demonstrate that our model achieves state-of-the-art performance on human-centric scene reconstruction and delivers highly competitive results on global human motion estimation, comparing favorably against both optimization-based frameworks and HMR-only methods. Project page: https://murphylmf.github.io/UniSH/
title UniSH: Unifying Scene and Human Reconstruction in a Feed-Forward Pass
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2601.01222