OpenHuman4D: Open-Vocabulary 4D Human Parsing

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Suzuki, Keito, Du, Bang, Li, Runfa Blark, Chen, Kunyao, Wang, Lei, Liu, Peng, Bi, Ning, Nguyen, Truong
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866909706028580864
author Suzuki, Keito
Du, Bang
Li, Runfa Blark
Chen, Kunyao
Wang, Lei
Liu, Peng
Bi, Ning
Nguyen, Truong
author_facet Suzuki, Keito
Du, Bang
Li, Runfa Blark
Chen, Kunyao
Wang, Lei
Liu, Peng
Bi, Ning
Nguyen, Truong
contents Understanding dynamic 3D human representation has become increasingly critical in virtual and extended reality applications. However, existing human part segmentation methods are constrained by reliance on closed-set datasets and prolonged inference times, which significantly restrict their applicability. In this paper, we introduce the first 4D human parsing framework that simultaneously addresses these challenges by reducing the inference time and introducing open-vocabulary capabilities. Building upon state-of-the-art open-vocabulary 3D human parsing techniques, our approach extends the support to 4D human-centric video with three key innovations: 1) We adopt mask-based video object tracking to efficiently establish spatial and temporal correspondences, avoiding the necessity of segmenting all frames. 2) A novel Mask Validation module is designed to manage new target identification and mitigate tracking failures. 3) We propose a 4D Mask Fusion module, integrating memory-conditioned attention and logits equalization for robust embedding fusion. Extensive experiments demonstrate the effectiveness and flexibility of the proposed method on 4D human-centric parsing tasks, achieving up to 93.3% acceleration compared to the previous state-of-the-art method, which was limited to parsing fixed classes.
format Preprint
id arxiv_https___arxiv_org_abs_2507_09880
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle OpenHuman4D: Open-Vocabulary 4D Human Parsing
Suzuki, Keito
Du, Bang
Li, Runfa Blark
Chen, Kunyao
Wang, Lei
Liu, Peng
Bi, Ning
Nguyen, Truong
Computer Vision and Pattern Recognition
Understanding dynamic 3D human representation has become increasingly critical in virtual and extended reality applications. However, existing human part segmentation methods are constrained by reliance on closed-set datasets and prolonged inference times, which significantly restrict their applicability. In this paper, we introduce the first 4D human parsing framework that simultaneously addresses these challenges by reducing the inference time and introducing open-vocabulary capabilities. Building upon state-of-the-art open-vocabulary 3D human parsing techniques, our approach extends the support to 4D human-centric video with three key innovations: 1) We adopt mask-based video object tracking to efficiently establish spatial and temporal correspondences, avoiding the necessity of segmenting all frames. 2) A novel Mask Validation module is designed to manage new target identification and mitigate tracking failures. 3) We propose a 4D Mask Fusion module, integrating memory-conditioned attention and logits equalization for robust embedding fusion. Extensive experiments demonstrate the effectiveness and flexibility of the proposed method on 4D human-centric parsing tasks, achieving up to 93.3% acceleration compared to the previous state-of-the-art method, which was limited to parsing fixed classes.
title OpenHuman4D: Open-Vocabulary 4D Human Parsing
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2507.09880