IPFormer: Visual 3D Panoptic Scene Completion with Context-Adaptive Instance Proposals

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gross, Markus, Fahmy, Aya, Niwattananan, Danit, Muhle, Dominik, Song, Rui, Cremers, Daniel, Meeß, Henri
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918303161647104
author Gross, Markus
Fahmy, Aya
Niwattananan, Danit
Muhle, Dominik
Song, Rui
Cremers, Daniel
Meeß, Henri
author_facet Gross, Markus
Fahmy, Aya
Niwattananan, Danit
Muhle, Dominik
Song, Rui
Cremers, Daniel
Meeß, Henri
contents Semantic Scene Completion (SSC) has emerged as a pivotal approach for jointly learning scene geometry and semantics, enabling downstream applications such as navigation in mobile robotics. The recent generalization to Panoptic Scene Completion (PSC) advances the SSC domain by integrating instance-level information, thereby enhancing object-level sensitivity in scene understanding. While PSC was introduced using LiDAR modality, methods based on camera images remain largely unexplored. Moreover, recent Transformer-based approaches utilize a fixed set of learned queries to reconstruct objects within the scene volume. Although these queries are typically updated with image context during training, they remain static at test time, limiting their ability to dynamically adapt specifically to the observed scene. To overcome these limitations, we propose IPFormer, the first method that leverages context-adaptive instance proposals at train and test time to address vision-based 3D Panoptic Scene Completion. Specifically, IPFormer adaptively initializes these queries as panoptic instance proposals derived from image context and further refines them through attention-based encoding and decoding to reason about semantic instance-voxel relationships. Extensive experimental results show that our approach achieves state-of-the-art in-domain performance, exhibits superior zero-shot generalization on out-of-domain data, and achieves a runtime reduction exceeding 14x. These results highlight our introduction of context-adaptive instance proposals as a pioneering effort in addressing vision-based 3D Panoptic Scene Completion. Code available at https://github.com/markus-42/ipformer.
format Preprint
id arxiv_https___arxiv_org_abs_2506_20671
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle IPFormer: Visual 3D Panoptic Scene Completion with Context-Adaptive Instance Proposals
Gross, Markus
Fahmy, Aya
Niwattananan, Danit
Muhle, Dominik
Song, Rui
Cremers, Daniel
Meeß, Henri
Computer Vision and Pattern Recognition
Semantic Scene Completion (SSC) has emerged as a pivotal approach for jointly learning scene geometry and semantics, enabling downstream applications such as navigation in mobile robotics. The recent generalization to Panoptic Scene Completion (PSC) advances the SSC domain by integrating instance-level information, thereby enhancing object-level sensitivity in scene understanding. While PSC was introduced using LiDAR modality, methods based on camera images remain largely unexplored. Moreover, recent Transformer-based approaches utilize a fixed set of learned queries to reconstruct objects within the scene volume. Although these queries are typically updated with image context during training, they remain static at test time, limiting their ability to dynamically adapt specifically to the observed scene. To overcome these limitations, we propose IPFormer, the first method that leverages context-adaptive instance proposals at train and test time to address vision-based 3D Panoptic Scene Completion. Specifically, IPFormer adaptively initializes these queries as panoptic instance proposals derived from image context and further refines them through attention-based encoding and decoding to reason about semantic instance-voxel relationships. Extensive experimental results show that our approach achieves state-of-the-art in-domain performance, exhibits superior zero-shot generalization on out-of-domain data, and achieves a runtime reduction exceeding 14x. These results highlight our introduction of context-adaptive instance proposals as a pioneering effort in addressing vision-based 3D Panoptic Scene Completion. Code available at https://github.com/markus-42/ipformer.
title IPFormer: Visual 3D Panoptic Scene Completion with Context-Adaptive Instance Proposals
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.20671