Segment then Splat: Unified 3D Open-Vocabulary Segmentation via Gaussian Splatting

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lu, Yiren, Zhou, Yunlai, Qiao, Yiran, Song, Chaoda, Liang, Tuo, Ma, Jing, Wang, Huan, Yin, Yu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915578255507456
author Lu, Yiren
Zhou, Yunlai
Qiao, Yiran
Song, Chaoda
Liang, Tuo
Ma, Jing
Wang, Huan
Yin, Yu
author_facet Lu, Yiren
Zhou, Yunlai
Qiao, Yiran
Song, Chaoda
Liang, Tuo
Ma, Jing
Wang, Huan
Yin, Yu
contents Open-vocabulary querying in 3D space is crucial for enabling more intelligent perception in applications such as robotics, autonomous systems, and augmented reality. However, most existing methods rely on 2D pixel-level parsing, leading to multi-view inconsistencies and poor 3D object retrieval. Moreover, they are limited to static scenes and struggle with dynamic scenes due to the complexities of motion modeling. In this paper, we propose Segment then Splat, a 3D-aware open vocabulary segmentation approach for both static and dynamic scenes based on Gaussian Splatting. Segment then Splat reverses the long established approach of "segmentation after reconstruction" by dividing Gaussians into distinct object sets before reconstruction. Once reconstruction is complete, the scene is naturally segmented into individual objects, achieving true 3D segmentation. This design eliminates both geometric and semantic ambiguities, as well as Gaussian-object misalignment issues in dynamic scenes. It also accelerates the optimization process, as it eliminates the need for learning a separate language field. After optimization, a CLIP embedding is assigned to each object to enable open-vocabulary querying. Extensive experiments one various datasets demonstrate the effectiveness of our proposed method in both static and dynamic scenarios.
format Preprint
id arxiv_https___arxiv_org_abs_2503_22204
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Segment then Splat: Unified 3D Open-Vocabulary Segmentation via Gaussian Splatting
Lu, Yiren
Zhou, Yunlai
Qiao, Yiran
Song, Chaoda
Liang, Tuo
Ma, Jing
Wang, Huan
Yin, Yu
Computer Vision and Pattern Recognition
Open-vocabulary querying in 3D space is crucial for enabling more intelligent perception in applications such as robotics, autonomous systems, and augmented reality. However, most existing methods rely on 2D pixel-level parsing, leading to multi-view inconsistencies and poor 3D object retrieval. Moreover, they are limited to static scenes and struggle with dynamic scenes due to the complexities of motion modeling. In this paper, we propose Segment then Splat, a 3D-aware open vocabulary segmentation approach for both static and dynamic scenes based on Gaussian Splatting. Segment then Splat reverses the long established approach of "segmentation after reconstruction" by dividing Gaussians into distinct object sets before reconstruction. Once reconstruction is complete, the scene is naturally segmented into individual objects, achieving true 3D segmentation. This design eliminates both geometric and semantic ambiguities, as well as Gaussian-object misalignment issues in dynamic scenes. It also accelerates the optimization process, as it eliminates the need for learning a separate language field. After optimization, a CLIP embedding is assigned to each object to enable open-vocabulary querying. Extensive experiments one various datasets demonstrate the effectiveness of our proposed method in both static and dynamic scenarios.
title Segment then Splat: Unified 3D Open-Vocabulary Segmentation via Gaussian Splatting
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.22204