OpenGaFF: Open-Vocabulary Gaussian Feature Field with Codebook Attention

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Kunyi, Niemeyer, Michael, Wang, Sen, Gasperini, Stefano, Navab, Nassir, Tombari, Federico
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914589730406400
author Li, Kunyi
Niemeyer, Michael
Wang, Sen
Gasperini, Stefano
Navab, Nassir
Tombari, Federico
author_facet Li, Kunyi
Niemeyer, Michael
Wang, Sen
Gasperini, Stefano
Navab, Nassir
Tombari, Federico
contents Understanding open-vocabulary 3D scenes with Gaussian-based representations remains challenging due to fragmented and spatially inconsistent semantic predictions across multi-view observations. In this paper, we present OpenGaFF, a novel framework for open-vocabulary 3D scene understanding built upon 3D Gaussian Splatting. At the core of our method is a Gaussian Feature Field that models semantics as a continuous function of Gaussian geometry and appearance. By explicitly conditioning semantic predictions on geometric structure, this formulation strengthens the coupling between geometry and semantics, leading to improved spatial coherence across similar structures in 3D space. To further enforce object-level semantic consistency, we introduce a structured codebook that serves as a set of shared semantic primitives. Furthermore, a codebook-guided attention mechanism is proposed to retrieve language features via similarity matching between query embeddings and learned codebook entries, enabling robust open-vocabulary reasoning while reducing intra-object feature variance. Extensive experiments on standard 2D and 3D open-vocabulary benchmarks demonstrate that our method consistently outperforms prior approaches, achieving improved segmentation quality, stronger 3D semantic consistency and a semantically interpretable codebook that provides insight into the learned representation.
format Preprint
id arxiv_https___arxiv_org_abs_2605_06088
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle OpenGaFF: Open-Vocabulary Gaussian Feature Field with Codebook Attention
Li, Kunyi
Niemeyer, Michael
Wang, Sen
Gasperini, Stefano
Navab, Nassir
Tombari, Federico
Computer Vision and Pattern Recognition
Understanding open-vocabulary 3D scenes with Gaussian-based representations remains challenging due to fragmented and spatially inconsistent semantic predictions across multi-view observations. In this paper, we present OpenGaFF, a novel framework for open-vocabulary 3D scene understanding built upon 3D Gaussian Splatting. At the core of our method is a Gaussian Feature Field that models semantics as a continuous function of Gaussian geometry and appearance. By explicitly conditioning semantic predictions on geometric structure, this formulation strengthens the coupling between geometry and semantics, leading to improved spatial coherence across similar structures in 3D space. To further enforce object-level semantic consistency, we introduce a structured codebook that serves as a set of shared semantic primitives. Furthermore, a codebook-guided attention mechanism is proposed to retrieve language features via similarity matching between query embeddings and learned codebook entries, enabling robust open-vocabulary reasoning while reducing intra-object feature variance. Extensive experiments on standard 2D and 3D open-vocabulary benchmarks demonstrate that our method consistently outperforms prior approaches, achieving improved segmentation quality, stronger 3D semantic consistency and a semantically interpretable codebook that provides insight into the learned representation.
title OpenGaFF: Open-Vocabulary Gaussian Feature Field with Codebook Attention
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.06088