PanoGS: Gaussian-based Panoptic Segmentation for 3D Open Vocabulary Scene Understanding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhai, Hongjia, Li, Hai, Li, Zhenzhe, Pan, Xiaokun, He, Yijia, Zhang, Guofeng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917966594965504
author Zhai, Hongjia
Li, Hai
Li, Zhenzhe
Pan, Xiaokun
He, Yijia
Zhang, Guofeng
author_facet Zhai, Hongjia
Li, Hai
Li, Zhenzhe
Pan, Xiaokun
He, Yijia
Zhang, Guofeng
contents Recently, 3D Gaussian Splatting (3DGS) has shown encouraging performance for open vocabulary scene understanding tasks. However, previous methods cannot distinguish 3D instance-level information, which usually predicts a heatmap between the scene feature and text query. In this paper, we propose PanoGS, a novel and effective 3D panoptic open vocabulary scene understanding approach. Technically, to learn accurate 3D language features that can scale to large indoor scenarios, we adopt the pyramid tri-plane to model the latent continuous parametric feature space and use a 3D feature decoder to regress the multi-view fused 2D feature cloud. Besides, we propose language-guided graph cuts that synergistically leverage reconstructed geometry and learned language cues to group 3D Gaussian primitives into a set of super-primitives. To obtain 3D consistent instance, we perform graph clustering based segmentation with SAM-guided edge affinity computation between different super-primitives. Extensive experiments on widely used datasets show better or more competitive performance on 3D panoptic open vocabulary scene understanding. Project page: \href{https://zju3dv.github.io/panogs}{https://zju3dv.github.io/panogs}.
format Preprint
id arxiv_https___arxiv_org_abs_2503_18107
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle PanoGS: Gaussian-based Panoptic Segmentation for 3D Open Vocabulary Scene Understanding
Zhai, Hongjia
Li, Hai
Li, Zhenzhe
Pan, Xiaokun
He, Yijia
Zhang, Guofeng
Computer Vision and Pattern Recognition
Recently, 3D Gaussian Splatting (3DGS) has shown encouraging performance for open vocabulary scene understanding tasks. However, previous methods cannot distinguish 3D instance-level information, which usually predicts a heatmap between the scene feature and text query. In this paper, we propose PanoGS, a novel and effective 3D panoptic open vocabulary scene understanding approach. Technically, to learn accurate 3D language features that can scale to large indoor scenarios, we adopt the pyramid tri-plane to model the latent continuous parametric feature space and use a 3D feature decoder to regress the multi-view fused 2D feature cloud. Besides, we propose language-guided graph cuts that synergistically leverage reconstructed geometry and learned language cues to group 3D Gaussian primitives into a set of super-primitives. To obtain 3D consistent instance, we perform graph clustering based segmentation with SAM-guided edge affinity computation between different super-primitives. Extensive experiments on widely used datasets show better or more competitive performance on 3D panoptic open vocabulary scene understanding. Project page: \href{https://zju3dv.github.io/panogs}{https://zju3dv.github.io/panogs}.
title PanoGS: Gaussian-based Panoptic Segmentation for 3D Open Vocabulary Scene Understanding
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.18107