C3G: Learning Compact 3D Representations with 2K Gaussians

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: An, Honggyu, Jung, Jaewoo, Kim, Mungyeom, Kim, Chaehyun, Jeon, Minkyeong, Han, Jisang, Fukuda, Kazumi, Narihira, Takuya, Ko, Hyuna, Kim, Junsu, Hong, Sunghwan, Mitsufuji, Yuki, Kim, Seungryong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908996511727616
author An, Honggyu
Jung, Jaewoo
Kim, Mungyeom
Kim, Chaehyun
Jeon, Minkyeong
Han, Jisang
Fukuda, Kazumi
Narihira, Takuya
Ko, Hyuna
Kim, Junsu
Hong, Sunghwan
Mitsufuji, Yuki
Kim, Seungryong
author_facet An, Honggyu
Jung, Jaewoo
Kim, Mungyeom
Kim, Chaehyun
Jeon, Minkyeong
Han, Jisang
Fukuda, Kazumi
Narihira, Takuya
Ko, Hyuna
Kim, Junsu
Hong, Sunghwan
Mitsufuji, Yuki
Kim, Seungryong
contents Reconstructing and understanding 3D scenes from unposed sparse views in a feed-forward manner remains as a challenging task in 3D computer vision. Recent approaches use per-pixel 3D Gaussian Splatting for reconstruction, followed by a 2D-to-3D feature lifting stage for scene understanding. However, they generate excessive redundant Gaussians, causing high memory overhead and sub-optimal multi-view feature aggregation, leading to degraded novel view synthesis and scene understanding performance. We propose C3G, a novel feed-forward framework that estimates compact 3D Gaussians only at essential spatial locations, minimizing redundancy while enabling effective feature lifting. We introduce learnable tokens that aggregate multi-view features through self-attention to guide Gaussian generation, ensuring each Gaussian integrates relevant visual features across views. We then exploit the learned attention patterns for Gaussian decoding to efficiently lift features. Extensive experiments on pose-free novel view synthesis, 3D open-vocabulary segmentation, and view-invariant feature aggregation demonstrate our approach's effectiveness. Results show that a compact yet geometrically meaningful representation is sufficient for high-quality scene reconstruction and understanding, achieving superior memory efficiency and feature fidelity compared to existing methods.
format Preprint
id arxiv_https___arxiv_org_abs_2512_04021
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle C3G: Learning Compact 3D Representations with 2K Gaussians
An, Honggyu
Jung, Jaewoo
Kim, Mungyeom
Kim, Chaehyun
Jeon, Minkyeong
Han, Jisang
Fukuda, Kazumi
Narihira, Takuya
Ko, Hyuna
Kim, Junsu
Hong, Sunghwan
Mitsufuji, Yuki
Kim, Seungryong
Computer Vision and Pattern Recognition
Reconstructing and understanding 3D scenes from unposed sparse views in a feed-forward manner remains as a challenging task in 3D computer vision. Recent approaches use per-pixel 3D Gaussian Splatting for reconstruction, followed by a 2D-to-3D feature lifting stage for scene understanding. However, they generate excessive redundant Gaussians, causing high memory overhead and sub-optimal multi-view feature aggregation, leading to degraded novel view synthesis and scene understanding performance. We propose C3G, a novel feed-forward framework that estimates compact 3D Gaussians only at essential spatial locations, minimizing redundancy while enabling effective feature lifting. We introduce learnable tokens that aggregate multi-view features through self-attention to guide Gaussian generation, ensuring each Gaussian integrates relevant visual features across views. We then exploit the learned attention patterns for Gaussian decoding to efficiently lift features. Extensive experiments on pose-free novel view synthesis, 3D open-vocabulary segmentation, and view-invariant feature aggregation demonstrate our approach's effectiveness. Results show that a compact yet geometrically meaningful representation is sufficient for high-quality scene reconstruction and understanding, achieving superior memory efficiency and feature fidelity compared to existing methods.
title C3G: Learning Compact 3D Representations with 2K Gaussians
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.04021