Towards 3D Object-Centric Feature Learning for Semantic Scene Completion

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Wang, Weihua, Cui, Yubo, Lin, Xiangru, Li, Zhiheng, Fang, Zheng
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866912781570146304
author Wang, Weihua
Cui, Yubo
Lin, Xiangru
Li, Zhiheng
Fang, Zheng
author_facet Wang, Weihua
Cui, Yubo
Lin, Xiangru
Li, Zhiheng
Fang, Zheng
contents Vision-based 3D Semantic Scene Completion (SSC) has received growing attention due to its potential in autonomous driving. While most existing approaches follow an ego-centric paradigm by aggregating and diffusing features over the entire scene, they often overlook fine-grained object-level details, leading to semantic and geometric ambiguities, especially in complex environments. To address this limitation, we propose Ocean, an object-centric prediction framework that decomposes the scene into individual object instances to enable more accurate semantic occupancy prediction. Specifically, we first employ a lightweight segmentation model, MobileSAM, to extract instance masks from the input image. Then, we introduce a 3D Semantic Group Attention module that leverages linear attention to aggregate object-centric features in 3D space. To handle segmentation errors and missing instances, we further design a Global Similarity-Guided Attention module that leverages segmentation features for global interaction. Finally, we propose an Instance-aware Local Diffusion module that improves instance features through a generative process and subsequently refines the scene representation in the BEV space. Extensive experiments on the SemanticKITTI and SSCBench-KITTI360 benchmarks demonstrate that Ocean achieves state-of-the-art performance, with mIoU scores of 17.40 and 20.28, respectively.
format Preprint
id arxiv_https___arxiv_org_abs_2511_13031
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Towards 3D Object-Centric Feature Learning for Semantic Scene Completion
Wang, Weihua
Cui, Yubo
Lin, Xiangru
Li, Zhiheng
Fang, Zheng
Computer Vision and Pattern Recognition
Vision-based 3D Semantic Scene Completion (SSC) has received growing attention due to its potential in autonomous driving. While most existing approaches follow an ego-centric paradigm by aggregating and diffusing features over the entire scene, they often overlook fine-grained object-level details, leading to semantic and geometric ambiguities, especially in complex environments. To address this limitation, we propose Ocean, an object-centric prediction framework that decomposes the scene into individual object instances to enable more accurate semantic occupancy prediction. Specifically, we first employ a lightweight segmentation model, MobileSAM, to extract instance masks from the input image. Then, we introduce a 3D Semantic Group Attention module that leverages linear attention to aggregate object-centric features in 3D space. To handle segmentation errors and missing instances, we further design a Global Similarity-Guided Attention module that leverages segmentation features for global interaction. Finally, we propose an Instance-aware Local Diffusion module that improves instance features through a generative process and subsequently refines the scene representation in the BEV space. Extensive experiments on the SemanticKITTI and SSCBench-KITTI360 benchmarks demonstrate that Ocean achieves state-of-the-art performance, with mIoU scores of 17.40 and 20.28, respectively.
title Towards 3D Object-Centric Feature Learning for Semantic Scene Completion
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.13031