Canonical Space Representation for 4D Panoptic Segmentation of Articulated Objects

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gomes, Manuel, Raducanu, Bogdan, Oliveira, Miguel
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915604611465216
author Gomes, Manuel
Raducanu, Bogdan
Oliveira, Miguel
author_facet Gomes, Manuel
Raducanu, Bogdan
Oliveira, Miguel
contents Articulated object perception presents significant challenges in computer vision, particularly because most existing methods ignore temporal dynamics despite the inherently dynamic nature of such objects. The use of 4D temporal data has not been thoroughly explored in articulated object perception and remains unexamined for panoptic segmentation. The lack of a benchmark dataset further hurt this field. To this end, we introduce Artic4D as a new dataset derived from PartNet Mobility and augmented with synthetic sensor data, featuring 4D panoptic annotations and articulation parameters. Building on this dataset, we propose CanonSeg4D, a novel 4D panoptic segmentation framework. This approach explicitly estimates per-frame offsets mapping observed object parts to a learned canonical space, thereby enhancing part-level segmentation. The framework employs this canonical representation to achieve consistent alignment of object parts across sequential frames. Comprehensive experiments on Artic4D demonstrate that the proposed CanonSeg4D outperforms state of the art approaches in panoptic segmentation accuracy in more complex scenarios. These findings highlight the effectiveness of temporal modeling and canonical alignment in dynamic object understanding, and pave the way for future advances in 4D articulated object perception.
format Preprint
id arxiv_https___arxiv_org_abs_2511_05356
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Canonical Space Representation for 4D Panoptic Segmentation of Articulated Objects
Gomes, Manuel
Raducanu, Bogdan
Oliveira, Miguel
Computer Vision and Pattern Recognition
I.2.10; I.4.6; I.5.1; I.5.4
Articulated object perception presents significant challenges in computer vision, particularly because most existing methods ignore temporal dynamics despite the inherently dynamic nature of such objects. The use of 4D temporal data has not been thoroughly explored in articulated object perception and remains unexamined for panoptic segmentation. The lack of a benchmark dataset further hurt this field. To this end, we introduce Artic4D as a new dataset derived from PartNet Mobility and augmented with synthetic sensor data, featuring 4D panoptic annotations and articulation parameters. Building on this dataset, we propose CanonSeg4D, a novel 4D panoptic segmentation framework. This approach explicitly estimates per-frame offsets mapping observed object parts to a learned canonical space, thereby enhancing part-level segmentation. The framework employs this canonical representation to achieve consistent alignment of object parts across sequential frames. Comprehensive experiments on Artic4D demonstrate that the proposed CanonSeg4D outperforms state of the art approaches in panoptic segmentation accuracy in more complex scenarios. These findings highlight the effectiveness of temporal modeling and canonical alignment in dynamic object understanding, and pave the way for future advances in 4D articulated object perception.
title Canonical Space Representation for 4D Panoptic Segmentation of Articulated Objects
topic Computer Vision and Pattern Recognition
I.2.10; I.4.6; I.5.1; I.5.4
url https://arxiv.org/abs/2511.05356