Scene-Agnostic Object-Centric Representation Learning for 3D Gaussian Splatting

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hsu, Tsuheng, Liu, Guiyu, Kannala, Juho, Heikkilä, Janne
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917398551986176
author Hsu, Tsuheng
Liu, Guiyu
Kannala, Juho
Heikkilä, Janne
author_facet Hsu, Tsuheng
Liu, Guiyu
Kannala, Juho
Heikkilä, Janne
contents Recent works on 3D scene understanding leverage 2D masks from visual foundation models (VFMs) to supervise radiance fields, enabling instance-level 3D segmentation. However, the supervision signals from foundation models are not fundamentally object-centric and often require additional mask pre/post-processing or specialized training and loss design to resolve mask identity conflicts across views. The learned identity of the 3D scene is scene-dependent, limiting generalizability across scenes. Therefore, we propose a dataset-level, object-centric supervision scheme to learn object representations in 3D Gaussian Splatting (3DGS). Building on a pre-trained slot attention-based Global Object Centric Learning (GOCL) module, we learn a scene-agnostic object codebook that provides consistent, identity-anchored representations across views and scenes. By coupling the codebook with the module's unsupervised object masks, we can directly supervise the identity features of 3D Gaussians without additional mask pre-/post-processing or explicit multi-view alignment. The learned scene-agnostic codebook enables object supervision and identification without per-scene fine-tuning or retraining. Our method thus introduces unsupervised object-centric learning (OCL) into 3DGS, yielding more structured representations and better generalization for downstream tasks such as robotic interaction, scene understanding, and cross-scene generalization.
format Preprint
id arxiv_https___arxiv_org_abs_2604_09045
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Scene-Agnostic Object-Centric Representation Learning for 3D Gaussian Splatting
Hsu, Tsuheng
Liu, Guiyu
Kannala, Juho
Heikkilä, Janne
Computer Vision and Pattern Recognition
Recent works on 3D scene understanding leverage 2D masks from visual foundation models (VFMs) to supervise radiance fields, enabling instance-level 3D segmentation. However, the supervision signals from foundation models are not fundamentally object-centric and often require additional mask pre/post-processing or specialized training and loss design to resolve mask identity conflicts across views. The learned identity of the 3D scene is scene-dependent, limiting generalizability across scenes. Therefore, we propose a dataset-level, object-centric supervision scheme to learn object representations in 3D Gaussian Splatting (3DGS). Building on a pre-trained slot attention-based Global Object Centric Learning (GOCL) module, we learn a scene-agnostic object codebook that provides consistent, identity-anchored representations across views and scenes. By coupling the codebook with the module's unsupervised object masks, we can directly supervise the identity features of 3D Gaussians without additional mask pre-/post-processing or explicit multi-view alignment. The learned scene-agnostic codebook enables object supervision and identification without per-scene fine-tuning or retraining. Our method thus introduces unsupervised object-centric learning (OCL) into 3DGS, yielding more structured representations and better generalization for downstream tasks such as robotic interaction, scene understanding, and cross-scene generalization.
title Scene-Agnostic Object-Centric Representation Learning for 3D Gaussian Splatting
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2604.09045