DCHM: Depth-Consistent Human Modeling for Multiview Detection

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Ma, Jiahao, Wang, Tianyu, Liu, Miaomiao, Ahmedt-Aristizabal, David, Nguyen, Chuong
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866918098460737536
author Ma, Jiahao
Wang, Tianyu
Liu, Miaomiao
Ahmedt-Aristizabal, David
Nguyen, Chuong
author_facet Ma, Jiahao
Wang, Tianyu
Liu, Miaomiao
Ahmedt-Aristizabal, David
Nguyen, Chuong
contents Multiview pedestrian detection typically involves two stages: human modeling and pedestrian localization. Human modeling represents pedestrians in 3D space by fusing multiview information, making its quality crucial for detection accuracy. However, existing methods often introduce noise and have low precision. While some approaches reduce noise by fitting on costly multiview 3D annotations, they often struggle to generalize across diverse scenes. To eliminate reliance on human-labeled annotations and accurately model humans, we propose Depth-Consistent Human Modeling (DCHM), a framework designed for consistent depth estimation and multiview fusion in global coordinates. Specifically, our proposed pipeline with superpixel-wise Gaussian Splatting achieves multiview depth consistency in sparse-view, large-scaled, and crowded scenarios, producing precise point clouds for pedestrian localization. Extensive validations demonstrate that our method significantly reduces noise during human modeling, outperforming previous state-of-the-art baselines. Additionally, to our knowledge, DCHM is the first to reconstruct pedestrians and perform multiview segmentation in such a challenging setting. Code is available on the \href{https://jiahao-ma.github.io/DCHM/}{project page}.
format Preprint
id arxiv_https___arxiv_org_abs_2507_14505
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DCHM: Depth-Consistent Human Modeling for Multiview Detection
Ma, Jiahao
Wang, Tianyu
Liu, Miaomiao
Ahmedt-Aristizabal, David
Nguyen, Chuong
Computer Vision and Pattern Recognition
Multiview pedestrian detection typically involves two stages: human modeling and pedestrian localization. Human modeling represents pedestrians in 3D space by fusing multiview information, making its quality crucial for detection accuracy. However, existing methods often introduce noise and have low precision. While some approaches reduce noise by fitting on costly multiview 3D annotations, they often struggle to generalize across diverse scenes. To eliminate reliance on human-labeled annotations and accurately model humans, we propose Depth-Consistent Human Modeling (DCHM), a framework designed for consistent depth estimation and multiview fusion in global coordinates. Specifically, our proposed pipeline with superpixel-wise Gaussian Splatting achieves multiview depth consistency in sparse-view, large-scaled, and crowded scenarios, producing precise point clouds for pedestrian localization. Extensive validations demonstrate that our method significantly reduces noise during human modeling, outperforming previous state-of-the-art baselines. Additionally, to our knowledge, DCHM is the first to reconstruct pedestrians and perform multiview segmentation in such a challenging setting. Code is available on the \href{https://jiahao-ma.github.io/DCHM/}{project page}.
title DCHM: Depth-Consistent Human Modeling for Multiview Detection
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2507.14505