GALA: Guided Attention with Language Alignment for Open Vocabulary Gaussian Splatting

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Alegret, Elena, Li, Kunyi, Wang, Sen, Liang, Siyun, Niemeyer, Michael, Gasperini, Stefano, Navab, Nassir, Tombari, Federico
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908496673374208
author Alegret, Elena
Li, Kunyi
Wang, Sen
Liang, Siyun
Niemeyer, Michael
Gasperini, Stefano
Navab, Nassir
Tombari, Federico
author_facet Alegret, Elena
Li, Kunyi
Wang, Sen
Liang, Siyun
Niemeyer, Michael
Gasperini, Stefano
Navab, Nassir
Tombari, Federico
contents 3D scene reconstruction and understanding have gained increasing popularity, yet existing methods still struggle to capture fine-grained, language-aware 3D representations from 2D images. In this paper, we present GALA, a novel framework for open-vocabulary 3D scene understanding with 3D Gaussian Splatting (3DGS). GALA distills a scene-specific 3D instance feature field via self-supervised contrastive learning. To extend to generalized language feature fields, we introduce the core contribution of GALA, a cross-attention module with two learnable codebooks that encode view-independent semantic embeddings. This design not only ensures intra-instance feature similarity but also supports seamless 2D and 3D open-vocabulary queries. It reduces memory consumption by avoiding per-Gaussian high-dimensional feature learning. Extensive experiments on real-world datasets demonstrate GALA's remarkable open-vocabulary performance on both 2D and 3D.
format Preprint
id arxiv_https___arxiv_org_abs_2508_14278
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle GALA: Guided Attention with Language Alignment for Open Vocabulary Gaussian Splatting
Alegret, Elena
Li, Kunyi
Wang, Sen
Liang, Siyun
Niemeyer, Michael
Gasperini, Stefano
Navab, Nassir
Tombari, Federico
Computer Vision and Pattern Recognition
3D scene reconstruction and understanding have gained increasing popularity, yet existing methods still struggle to capture fine-grained, language-aware 3D representations from 2D images. In this paper, we present GALA, a novel framework for open-vocabulary 3D scene understanding with 3D Gaussian Splatting (3DGS). GALA distills a scene-specific 3D instance feature field via self-supervised contrastive learning. To extend to generalized language feature fields, we introduce the core contribution of GALA, a cross-attention module with two learnable codebooks that encode view-independent semantic embeddings. This design not only ensures intra-instance feature similarity but also supports seamless 2D and 3D open-vocabulary queries. It reduces memory consumption by avoiding per-Gaussian high-dimensional feature learning. Extensive experiments on real-world datasets demonstrate GALA's remarkable open-vocabulary performance on both 2D and 3D.
title GALA: Guided Attention with Language Alignment for Open Vocabulary Gaussian Splatting
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2508.14278