Leveraging Multi-View Weak Supervision for Occlusion-Aware Multi-Human Parsing

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Bragagnolo, Laura, Terreran, Matteo, Barcellona, Leonardo, Ghidoni, Stefano
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908535343808512
author Bragagnolo, Laura
Terreran, Matteo
Barcellona, Leonardo
Ghidoni, Stefano
author_facet Bragagnolo, Laura
Terreran, Matteo
Barcellona, Leonardo
Ghidoni, Stefano
contents Multi-human parsing is the task of segmenting human body parts while associating each part to the person it belongs to, combining instance-level and part-level information for fine-grained human understanding. In this work, we demonstrate that, while state-of-the-art approaches achieved notable results on public datasets, they struggle considerably in segmenting people with overlapping bodies. From the intuition that overlapping people may appear separated from a different point of view, we propose a novel training framework exploiting multi-view information to improve multi-human parsing models under occlusions. Our method integrates such knowledge during the training process, introducing a novel approach based on weak supervision on human instances and a multi-view consistency loss. Given the lack of suitable datasets in the literature, we propose a semi-automatic annotation strategy to generate human instance segmentation masks from multi-view RGB+D data and 3D human skeletons. The experiments demonstrate that the approach can achieve up to a 4.20\% relative improvement on human parsing over the baseline model in occlusion scenarios.
format Preprint
id arxiv_https___arxiv_org_abs_2509_10093
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Leveraging Multi-View Weak Supervision for Occlusion-Aware Multi-Human Parsing
Bragagnolo, Laura
Terreran, Matteo
Barcellona, Leonardo
Ghidoni, Stefano
Computer Vision and Pattern Recognition
Multi-human parsing is the task of segmenting human body parts while associating each part to the person it belongs to, combining instance-level and part-level information for fine-grained human understanding. In this work, we demonstrate that, while state-of-the-art approaches achieved notable results on public datasets, they struggle considerably in segmenting people with overlapping bodies. From the intuition that overlapping people may appear separated from a different point of view, we propose a novel training framework exploiting multi-view information to improve multi-human parsing models under occlusions. Our method integrates such knowledge during the training process, introducing a novel approach based on weak supervision on human instances and a multi-view consistency loss. Given the lack of suitable datasets in the literature, we propose a semi-automatic annotation strategy to generate human instance segmentation masks from multi-view RGB+D data and 3D human skeletons. The experiments demonstrate that the approach can achieve up to a 4.20\% relative improvement on human parsing over the baseline model in occlusion scenarios.
title Leveraging Multi-View Weak Supervision for Occlusion-Aware Multi-Human Parsing
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2509.10093