GeoGuide: Hierarchical Geometric Guidance for Open-Vocabulary 3D Semantic Segmentation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tao, Xujing, Wang, Chuxin, Ai, Yubo, Cheng, Zhixin, Li, Zhuoyuan, Liu, Liangsheng, Chen, Yujia, Li, Xinjun, Li, Qiao, Yang, Wenfei, Zhang, Tianzhu
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908916920614912
author Tao, Xujing
Wang, Chuxin
Ai, Yubo
Cheng, Zhixin
Li, Zhuoyuan
Liu, Liangsheng
Chen, Yujia
Li, Xinjun
Li, Qiao
Yang, Wenfei
Zhang, Tianzhu
author_facet Tao, Xujing
Wang, Chuxin
Ai, Yubo
Cheng, Zhixin
Li, Zhuoyuan
Liu, Liangsheng
Chen, Yujia
Li, Xinjun
Li, Qiao
Yang, Wenfei
Zhang, Tianzhu
contents Open-vocabulary 3D semantic segmentation aims to segment arbitrary categories beyond the training set. Existing methods predominantly rely on distilling knowledge from 2D open-vocabulary models. However, aligning 3D features to the 2D representation space restricts intrinsic 3D geometric learning and inherits errors from 2D predictions. To address these limitations, we propose GeoGuide, a novel framework that leverages pretrained 3D models to integrate hierarchical geometry-semantic consistency for open-vocabulary 3D segmentation. Specifically, we introduce an Uncertainty-based Superpoint Distillation module to fuse geometric and semantic features for estimating per-point uncertainty, adaptively weighting 2D features within superpoints to suppress noise while preserving discriminative information to enhance local semantic consistency. Furthermore, our Instance-level Mask Reconstruction module leverages geometric priors to enforce semantic consistency within instances by reconstructing complete instance masks. Additionally, our Inter-Instance Relation Consistency module aligns geometric and semantic similarity matrices to calibrate cross-instance consistency for same-category objects, mitigating viewpoint-induced semantic drift. Extensive experiments on ScanNet v2, Matterport3D, and nuScenes demonstrate the superior performance of GeoGuide.
format Preprint
id arxiv_https___arxiv_org_abs_2603_26260
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle GeoGuide: Hierarchical Geometric Guidance for Open-Vocabulary 3D Semantic Segmentation
Tao, Xujing
Wang, Chuxin
Ai, Yubo
Cheng, Zhixin
Li, Zhuoyuan
Liu, Liangsheng
Chen, Yujia
Li, Xinjun
Li, Qiao
Yang, Wenfei
Zhang, Tianzhu
Computer Vision and Pattern Recognition
Artificial Intelligence
Open-vocabulary 3D semantic segmentation aims to segment arbitrary categories beyond the training set. Existing methods predominantly rely on distilling knowledge from 2D open-vocabulary models. However, aligning 3D features to the 2D representation space restricts intrinsic 3D geometric learning and inherits errors from 2D predictions. To address these limitations, we propose GeoGuide, a novel framework that leverages pretrained 3D models to integrate hierarchical geometry-semantic consistency for open-vocabulary 3D segmentation. Specifically, we introduce an Uncertainty-based Superpoint Distillation module to fuse geometric and semantic features for estimating per-point uncertainty, adaptively weighting 2D features within superpoints to suppress noise while preserving discriminative information to enhance local semantic consistency. Furthermore, our Instance-level Mask Reconstruction module leverages geometric priors to enforce semantic consistency within instances by reconstructing complete instance masks. Additionally, our Inter-Instance Relation Consistency module aligns geometric and semantic similarity matrices to calibrate cross-instance consistency for same-category objects, mitigating viewpoint-induced semantic drift. Extensive experiments on ScanNet v2, Matterport3D, and nuScenes demonstrate the superior performance of GeoGuide.
title GeoGuide: Hierarchical Geometric Guidance for Open-Vocabulary 3D Semantic Segmentation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2603.26260