Monocular Models are Strong Learners for Multi-View Human Mesh Recovery
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866910090950344704 |
|---|---|
| author | Xie, Haoyu Xu, Shengkai Guo, Cheng Saleem, Muhammad Usama Wu, Wenhan Chen, Chen Helmy, Ahmed Wang, Pu Xue, Hongfei |
| author_facet | Xie, Haoyu Xu, Shengkai Guo, Cheng Saleem, Muhammad Usama Wu, Wenhan Chen, Chen Helmy, Ahmed Wang, Pu Xue, Hongfei |
| contents | Multi-view human mesh recovery (HMR) is broadly deployed in diverse domains where high accuracy and strong generalization are essential. Existing approaches can be broadly grouped into geometry-based and learning-based methods. However, geometry-based methods (e.g., triangulation) rely on cumbersome camera calibration, while learning-based approaches often generalize poorly to unseen camera configurations due to the lack of multi-view training data, limiting their performance in real-world scenarios. To enable calibration-free reconstruction that generalizes to arbitrary camera setups, we propose a training-free framework that leverages pretrained single-view HMR models as strong priors, eliminating the need for multi-view training data. Our method first constructs a robust and consistent multi-view initialization from single-view predictions, and then refines it via test-time optimization guided by multi-view consistency and anatomical constraints. Extensive experiments demonstrate state-of-the-art performance on standard benchmarks, surpassing multi-view models trained with explicit multi-view supervision. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2603_20391 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Monocular Models are Strong Learners for Multi-View Human Mesh Recovery Xie, Haoyu Xu, Shengkai Guo, Cheng Saleem, Muhammad Usama Wu, Wenhan Chen, Chen Helmy, Ahmed Wang, Pu Xue, Hongfei Computer Vision and Pattern Recognition Multi-view human mesh recovery (HMR) is broadly deployed in diverse domains where high accuracy and strong generalization are essential. Existing approaches can be broadly grouped into geometry-based and learning-based methods. However, geometry-based methods (e.g., triangulation) rely on cumbersome camera calibration, while learning-based approaches often generalize poorly to unseen camera configurations due to the lack of multi-view training data, limiting their performance in real-world scenarios. To enable calibration-free reconstruction that generalizes to arbitrary camera setups, we propose a training-free framework that leverages pretrained single-view HMR models as strong priors, eliminating the need for multi-view training data. Our method first constructs a robust and consistent multi-view initialization from single-view predictions, and then refines it via test-time optimization guided by multi-view consistency and anatomical constraints. Extensive experiments demonstrate state-of-the-art performance on standard benchmarks, surpassing multi-view models trained with explicit multi-view supervision. |
| title | Monocular Models are Strong Learners for Multi-View Human Mesh Recovery |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2603.20391 |