Feed-Forward SceneDINO for Unsupervised Semantic Scene Completion

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jevtić, Aleksandar, Reich, Christoph, Wimbauer, Felix, Hahn, Oliver, Rupprecht, Christian, Roth, Stefan, Cremers, Daniel
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915408815063040
author Jevtić, Aleksandar
Reich, Christoph
Wimbauer, Felix
Hahn, Oliver
Rupprecht, Christian
Roth, Stefan
Cremers, Daniel
author_facet Jevtić, Aleksandar
Reich, Christoph
Wimbauer, Felix
Hahn, Oliver
Rupprecht, Christian
Roth, Stefan
Cremers, Daniel
contents Semantic scene completion (SSC) aims to infer both the 3D geometry and semantics of a scene from single images. In contrast to prior work on SSC that heavily relies on expensive ground-truth annotations, we approach SSC in an unsupervised setting. Our novel method, SceneDINO, adapts techniques from self-supervised representation learning and 2D unsupervised scene understanding to SSC. Our training exclusively utilizes multi-view consistency self-supervision without any form of semantic or geometric ground truth. Given a single input image, SceneDINO infers the 3D geometry and expressive 3D DINO features in a feed-forward manner. Through a novel 3D feature distillation approach, we obtain unsupervised 3D semantics. In both 3D and 2D unsupervised scene understanding, SceneDINO reaches state-of-the-art segmentation accuracy. Linear probing our 3D features matches the segmentation accuracy of a current supervised SSC approach. Additionally, we showcase the domain generalization and multi-view consistency of SceneDINO, taking the first steps towards a strong foundation for single image 3D scene understanding.
format Preprint
id arxiv_https___arxiv_org_abs_2507_06230
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Feed-Forward SceneDINO for Unsupervised Semantic Scene Completion
Jevtić, Aleksandar
Reich, Christoph
Wimbauer, Felix
Hahn, Oliver
Rupprecht, Christian
Roth, Stefan
Cremers, Daniel
Computer Vision and Pattern Recognition
Semantic scene completion (SSC) aims to infer both the 3D geometry and semantics of a scene from single images. In contrast to prior work on SSC that heavily relies on expensive ground-truth annotations, we approach SSC in an unsupervised setting. Our novel method, SceneDINO, adapts techniques from self-supervised representation learning and 2D unsupervised scene understanding to SSC. Our training exclusively utilizes multi-view consistency self-supervision without any form of semantic or geometric ground truth. Given a single input image, SceneDINO infers the 3D geometry and expressive 3D DINO features in a feed-forward manner. Through a novel 3D feature distillation approach, we obtain unsupervised 3D semantics. In both 3D and 2D unsupervised scene understanding, SceneDINO reaches state-of-the-art segmentation accuracy. Linear probing our 3D features matches the segmentation accuracy of a current supervised SSC approach. Additionally, we showcase the domain generalization and multi-view consistency of SceneDINO, taking the first steps towards a strong foundation for single image 3D scene understanding.
title Feed-Forward SceneDINO for Unsupervised Semantic Scene Completion
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2507.06230