Monocular Models are Strong Learners for Multi-View Human Mesh Recovery

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xie, Haoyu, Xu, Shengkai, Guo, Cheng, Saleem, Muhammad Usama, Wu, Wenhan, Chen, Chen, Helmy, Ahmed, Wang, Pu, Xue, Hongfei
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910090950344704
author Xie, Haoyu
Xu, Shengkai
Guo, Cheng
Saleem, Muhammad Usama
Wu, Wenhan
Chen, Chen
Helmy, Ahmed
Wang, Pu
Xue, Hongfei
author_facet Xie, Haoyu
Xu, Shengkai
Guo, Cheng
Saleem, Muhammad Usama
Wu, Wenhan
Chen, Chen
Helmy, Ahmed
Wang, Pu
Xue, Hongfei
contents Multi-view human mesh recovery (HMR) is broadly deployed in diverse domains where high accuracy and strong generalization are essential. Existing approaches can be broadly grouped into geometry-based and learning-based methods. However, geometry-based methods (e.g., triangulation) rely on cumbersome camera calibration, while learning-based approaches often generalize poorly to unseen camera configurations due to the lack of multi-view training data, limiting their performance in real-world scenarios. To enable calibration-free reconstruction that generalizes to arbitrary camera setups, we propose a training-free framework that leverages pretrained single-view HMR models as strong priors, eliminating the need for multi-view training data. Our method first constructs a robust and consistent multi-view initialization from single-view predictions, and then refines it via test-time optimization guided by multi-view consistency and anatomical constraints. Extensive experiments demonstrate state-of-the-art performance on standard benchmarks, surpassing multi-view models trained with explicit multi-view supervision.
format Preprint
id arxiv_https___arxiv_org_abs_2603_20391
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Monocular Models are Strong Learners for Multi-View Human Mesh Recovery
Xie, Haoyu
Xu, Shengkai
Guo, Cheng
Saleem, Muhammad Usama
Wu, Wenhan
Chen, Chen
Helmy, Ahmed
Wang, Pu
Xue, Hongfei
Computer Vision and Pattern Recognition
Multi-view human mesh recovery (HMR) is broadly deployed in diverse domains where high accuracy and strong generalization are essential. Existing approaches can be broadly grouped into geometry-based and learning-based methods. However, geometry-based methods (e.g., triangulation) rely on cumbersome camera calibration, while learning-based approaches often generalize poorly to unseen camera configurations due to the lack of multi-view training data, limiting their performance in real-world scenarios. To enable calibration-free reconstruction that generalizes to arbitrary camera setups, we propose a training-free framework that leverages pretrained single-view HMR models as strong priors, eliminating the need for multi-view training data. Our method first constructs a robust and consistent multi-view initialization from single-view predictions, and then refines it via test-time optimization guided by multi-view consistency and anatomical constraints. Extensive experiments demonstrate state-of-the-art performance on standard benchmarks, surpassing multi-view models trained with explicit multi-view supervision.
title Monocular Models are Strong Learners for Multi-View Human Mesh Recovery
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2603.20391