LHM++: An Efficient Large Human Reconstruction Model for Pose-free Images to 3D

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Qiu, Lingteng, Li, Peihao, Li, Heyuan, Zuo, Qi, Gu, Xiaodong, Dong, Yuan, Yuan, Weihao, Peng, Rui, Zhu, Siyu, Han, Xiaoguang, Chen, Guanying, Dong, Zilong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914396817588224
author Qiu, Lingteng
Li, Peihao
Li, Heyuan
Zuo, Qi
Gu, Xiaodong
Dong, Yuan
Yuan, Weihao
Peng, Rui
Zhu, Siyu
Han, Xiaoguang
Chen, Guanying
Dong, Zilong
author_facet Qiu, Lingteng
Li, Peihao
Li, Heyuan
Zuo, Qi
Gu, Xiaodong
Dong, Yuan
Yuan, Weihao
Peng, Rui
Zhu, Siyu
Han, Xiaoguang
Chen, Guanying
Dong, Zilong
contents Reconstructing animatable 3D humans from casually captured images of articulated subjects without camera or pose information is highly practical but remains challenging due to view misalignment, occlusions, and the absence of structural priors. In this work, we present LHM++, an efficient large-scale human reconstruction model that generates high-quality, animatable 3D avatars within seconds from one or multiple pose-free images. At its core is an Encoder-Decoder Point-Image Transformer architecture that progressively encodes and decodes 3D geometric point features to improve efficiency, while fusing hierarchical 3D point features with image features through multimodal attention. The fused features are decoded into 3D Gaussian splats to recover detailed geometry and appearance. To further enhance visual fidelity, we introduce a lightweight 3D-aware neural animation renderer that refines the rendering quality of reconstructed avatars in real time. Extensive experiments show that our method produces high-fidelity, animatable 3D humans without requiring camera or pose annotations. Our code and project page are available at https://lingtengqiu.github.io/LHM++/
format Preprint
id arxiv_https___arxiv_org_abs_2506_13766
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LHM++: An Efficient Large Human Reconstruction Model for Pose-free Images to 3D
Qiu, Lingteng
Li, Peihao
Li, Heyuan
Zuo, Qi
Gu, Xiaodong
Dong, Yuan
Yuan, Weihao
Peng, Rui
Zhu, Siyu
Han, Xiaoguang
Chen, Guanying
Dong, Zilong
Computer Vision and Pattern Recognition
Reconstructing animatable 3D humans from casually captured images of articulated subjects without camera or pose information is highly practical but remains challenging due to view misalignment, occlusions, and the absence of structural priors. In this work, we present LHM++, an efficient large-scale human reconstruction model that generates high-quality, animatable 3D avatars within seconds from one or multiple pose-free images. At its core is an Encoder-Decoder Point-Image Transformer architecture that progressively encodes and decodes 3D geometric point features to improve efficiency, while fusing hierarchical 3D point features with image features through multimodal attention. The fused features are decoded into 3D Gaussian splats to recover detailed geometry and appearance. To further enhance visual fidelity, we introduce a lightweight 3D-aware neural animation renderer that refines the rendering quality of reconstructed avatars in real time. Extensive experiments show that our method produces high-fidelity, animatable 3D humans without requiring camera or pose annotations. Our code and project page are available at https://lingtengqiu.github.io/LHM++/
title LHM++: An Efficient Large Human Reconstruction Model for Pose-free Images to 3D
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.13766