SPNeRF: Open Vocabulary 3D Neural Scene Segmentation with Superpoints

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hu, Weiwen, Parodi, Niccolò, Zepp, Marcus, Feldmann, Ingo, Schreer, Oliver, Eisert, Peter
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915206308823040
author Hu, Weiwen
Parodi, Niccolò
Zepp, Marcus
Feldmann, Ingo
Schreer, Oliver
Eisert, Peter
author_facet Hu, Weiwen
Parodi, Niccolò
Zepp, Marcus
Feldmann, Ingo
Schreer, Oliver
Eisert, Peter
contents Open-vocabulary segmentation, powered by large visual-language models like CLIP, has expanded 2D segmentation capabilities beyond fixed classes predefined by the dataset, enabling zero-shot understanding across diverse scenes. Extending these capabilities to 3D segmentation introduces challenges, as CLIP's image-based embeddings often lack the geometric detail necessary for 3D scene segmentation. Recent methods tend to address this by introducing additional segmentation models or replacing CLIP with variations trained on segmentation data, which lead to redundancy or loss on CLIP's general language capabilities. To overcome this limitation, we introduce SPNeRF, a NeRF based zero-shot 3D segmentation approach that leverages geometric priors. We integrate geometric primitives derived from the 3D scene into NeRF training to produce primitive-wise CLIP features, avoiding the ambiguity of point-wise features. Additionally, we propose a primitive-based merging mechanism enhanced with affinity scores. Without relying on additional segmentation models, our method further explores CLIP's capability for 3D segmentation and achieves notable improvements over original LERF.
format Preprint
id arxiv_https___arxiv_org_abs_2503_15712
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SPNeRF: Open Vocabulary 3D Neural Scene Segmentation with Superpoints
Hu, Weiwen
Parodi, Niccolò
Zepp, Marcus
Feldmann, Ingo
Schreer, Oliver
Eisert, Peter
Computer Vision and Pattern Recognition
Open-vocabulary segmentation, powered by large visual-language models like CLIP, has expanded 2D segmentation capabilities beyond fixed classes predefined by the dataset, enabling zero-shot understanding across diverse scenes. Extending these capabilities to 3D segmentation introduces challenges, as CLIP's image-based embeddings often lack the geometric detail necessary for 3D scene segmentation. Recent methods tend to address this by introducing additional segmentation models or replacing CLIP with variations trained on segmentation data, which lead to redundancy or loss on CLIP's general language capabilities. To overcome this limitation, we introduce SPNeRF, a NeRF based zero-shot 3D segmentation approach that leverages geometric priors. We integrate geometric primitives derived from the 3D scene into NeRF training to produce primitive-wise CLIP features, avoiding the ambiguity of point-wise features. Additionally, we propose a primitive-based merging mechanism enhanced with affinity scores. Without relying on additional segmentation models, our method further explores CLIP's capability for 3D segmentation and achieves notable improvements over original LERF.
title SPNeRF: Open Vocabulary 3D Neural Scene Segmentation with Superpoints
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.15712