Coherent Human-Scene Reconstruction from Multi-Person Multi-View Video in a Single Pass

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kim, Sangmin, Hwang, Minhyuk, Cha, Geonho, Wee, Dongyoon, Park, Jaesik
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915870589059072
author Kim, Sangmin
Hwang, Minhyuk
Cha, Geonho
Wee, Dongyoon
Park, Jaesik
author_facet Kim, Sangmin
Hwang, Minhyuk
Cha, Geonho
Wee, Dongyoon
Park, Jaesik
contents Recent advances in 3D foundation models have led to growing interest in reconstructing humans and their surrounding environments. However, most existing approaches focus on monocular inputs, and extending them to multi-view settings requires additional overhead modules or preprocessed data. To this end, we present CHROMM, a unified framework that jointly estimates cameras, scene point clouds, and human meshes from multi-person multi-view videos without relying on external modules or preprocessing. We integrate strong geometric and human priors from Pi3X and Multi-HMR into a single trainable neural network architecture, and introduce a scale adjustment module to solve the scale discrepancy between humans and the scene. We also introduce a multi-view fusion strategy to aggregate per-view estimates into a single representation at test-time. Finally, we propose a geometry-based multi-person association method, which is more robust than appearance-based approaches. Experiments on EMDB, RICH, EgoHumans, and EgoExo4D show that CHROMM achieves competitive performance in global human motion and multi-view pose estimation while running over 8x faster than prior optimization-based multi-view approaches. Project page: https://nstar1125.github.io/chromm.
format Preprint
id arxiv_https___arxiv_org_abs_2603_12789
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Coherent Human-Scene Reconstruction from Multi-Person Multi-View Video in a Single Pass
Kim, Sangmin
Hwang, Minhyuk
Cha, Geonho
Wee, Dongyoon
Park, Jaesik
Computer Vision and Pattern Recognition
Recent advances in 3D foundation models have led to growing interest in reconstructing humans and their surrounding environments. However, most existing approaches focus on monocular inputs, and extending them to multi-view settings requires additional overhead modules or preprocessed data. To this end, we present CHROMM, a unified framework that jointly estimates cameras, scene point clouds, and human meshes from multi-person multi-view videos without relying on external modules or preprocessing. We integrate strong geometric and human priors from Pi3X and Multi-HMR into a single trainable neural network architecture, and introduce a scale adjustment module to solve the scale discrepancy between humans and the scene. We also introduce a multi-view fusion strategy to aggregate per-view estimates into a single representation at test-time. Finally, we propose a geometry-based multi-person association method, which is more robust than appearance-based approaches. Experiments on EMDB, RICH, EgoHumans, and EgoExo4D show that CHROMM achieves competitive performance in global human motion and multi-view pose estimation while running over 8x faster than prior optimization-based multi-view approaches. Project page: https://nstar1125.github.io/chromm.
title Coherent Human-Scene Reconstruction from Multi-Person Multi-View Video in a Single Pass
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2603.12789