From Camera to World: A Plug-and-Play Module for Human Mesh Transformation

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Ma, Changhai, Wu, Ziyu, Zhang, Yunkang, Ying, Qijun, Liu, Boyan, Cai, Xiaohui
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912770774007808
author Ma, Changhai
Wu, Ziyu
Zhang, Yunkang
Ying, Qijun
Liu, Boyan
Cai, Xiaohui
author_facet Ma, Changhai
Wu, Ziyu
Zhang, Yunkang
Ying, Qijun
Liu, Boyan
Cai, Xiaohui
contents Reconstructing accurate 3D human meshes in the world coordinate system from in-the-wild images remains challenging due to the lack of camera rotation information. While existing methods achieve promising results in the camera coordinate system by assuming zero camera rotation, this simplification leads to significant errors when transforming the reconstructed mesh to the world coordinate system. To address this challenge, we propose Mesh-Plug, a plug-and-play module that accurately transforms human meshes from camera coordinates to world coordinates. Our key innovation lies in a human-centered approach that leverages both RGB images and depth maps rendered from the initial mesh to estimate camera rotation parameters, eliminating the dependency on environmental cues. Specifically, we first train a camera rotation prediction module that focuses on the human body's spatial configuration to estimate camera pitch angle. Then, by integrating the predicted camera parameters with the initial mesh, we design a mesh adjustment module that simultaneously refines the root joint orientation and body pose. Extensive experiments demonstrate that our framework outperforms state-of-the-art methods on the benchmark datasets SPEC-SYN and SPEC-MTP.
format Preprint
id arxiv_https___arxiv_org_abs_2512_15212
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle From Camera to World: A Plug-and-Play Module for Human Mesh Transformation
Ma, Changhai
Wu, Ziyu
Zhang, Yunkang
Ying, Qijun
Liu, Boyan
Cai, Xiaohui
Computer Vision and Pattern Recognition
Reconstructing accurate 3D human meshes in the world coordinate system from in-the-wild images remains challenging due to the lack of camera rotation information. While existing methods achieve promising results in the camera coordinate system by assuming zero camera rotation, this simplification leads to significant errors when transforming the reconstructed mesh to the world coordinate system. To address this challenge, we propose Mesh-Plug, a plug-and-play module that accurately transforms human meshes from camera coordinates to world coordinates. Our key innovation lies in a human-centered approach that leverages both RGB images and depth maps rendered from the initial mesh to estimate camera rotation parameters, eliminating the dependency on environmental cues. Specifically, we first train a camera rotation prediction module that focuses on the human body's spatial configuration to estimate camera pitch angle. Then, by integrating the predicted camera parameters with the initial mesh, we design a mesh adjustment module that simultaneously refines the root joint orientation and body pose. Extensive experiments demonstrate that our framework outperforms state-of-the-art methods on the benchmark datasets SPEC-SYN and SPEC-MTP.
title From Camera to World: A Plug-and-Play Module for Human Mesh Transformation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.15212