Chorus: Multi-Teacher Pretraining for Holistic 3D Gaussian Scene Encoding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Yue, Ma, Qi, Yang, Runyi, Ma, Mengjiao, Ren, Bin, Popovic, Nikola, Sebe, Nicu, Gevers, Theo, Van Gool, Luc, Paudel, Danda Pani, Oswald, Martin R.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910190127808512
author Li, Yue
Ma, Qi
Yang, Runyi
Ma, Mengjiao
Ren, Bin
Popovic, Nikola
Sebe, Nicu
Gevers, Theo
Van Gool, Luc
Paudel, Danda Pani
Oswald, Martin R.
author_facet Li, Yue
Ma, Qi
Yang, Runyi
Ma, Mengjiao
Ren, Bin
Popovic, Nikola
Sebe, Nicu
Gevers, Theo
Van Gool, Luc
Paudel, Danda Pani
Oswald, Martin R.
contents While 3DGS has emerged as a high-fidelity scene representation, encoding rich, general-purpose features directly from its primitives remains under-explored. We address this gap by introducing Chorus, a multi-teacher pretraining framework that learns a holistic feed-forward 3D Gaussian Splatting (3DGS) scene encoder by distilling complementary signals from 2D foundation models. Chorus employs a shared 3D encoder and teacher-specific projectors to learn from language-aligned, generalist, and object-aware teachers, encouraging a shared embedding space that captures signals from high-level semantics to fine-grained structure. We evaluate Chorus on a wide range of tasks: open-vocabulary semantic and instance segmentation, linear and decoder probing, data-efficient supervision, as well as LLM-based Q&A. Besides 3DGS, we also test Chorus on several benchmarks that only support point clouds by pretraining a variant using only Gaussian centers, colors, and estimated normals. Surprisingly, this encoder shows strong transfer and outperforms the point-cloud baseline while using 39.9 times fewer training scenes. Finally, we propose a render-and-distill adaptation that facilitates out-of-domain finetuning.
format Preprint
id arxiv_https___arxiv_org_abs_2512_17817
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Chorus: Multi-Teacher Pretraining for Holistic 3D Gaussian Scene Encoding
Li, Yue
Ma, Qi
Yang, Runyi
Ma, Mengjiao
Ren, Bin
Popovic, Nikola
Sebe, Nicu
Gevers, Theo
Van Gool, Luc
Paudel, Danda Pani
Oswald, Martin R.
Computer Vision and Pattern Recognition
While 3DGS has emerged as a high-fidelity scene representation, encoding rich, general-purpose features directly from its primitives remains under-explored. We address this gap by introducing Chorus, a multi-teacher pretraining framework that learns a holistic feed-forward 3D Gaussian Splatting (3DGS) scene encoder by distilling complementary signals from 2D foundation models. Chorus employs a shared 3D encoder and teacher-specific projectors to learn from language-aligned, generalist, and object-aware teachers, encouraging a shared embedding space that captures signals from high-level semantics to fine-grained structure. We evaluate Chorus on a wide range of tasks: open-vocabulary semantic and instance segmentation, linear and decoder probing, data-efficient supervision, as well as LLM-based Q&A. Besides 3DGS, we also test Chorus on several benchmarks that only support point clouds by pretraining a variant using only Gaussian centers, colors, and estimated normals. Surprisingly, this encoder shows strong transfer and outperforms the point-cloud baseline while using 39.9 times fewer training scenes. Finally, we propose a render-and-distill adaptation that facilitates out-of-domain finetuning.
title Chorus: Multi-Teacher Pretraining for Holistic 3D Gaussian Scene Encoding
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.17817