OccLE: Label-Efficient 3D Semantic Occupancy Prediction

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Fang, Naiyu, Zhou, Zheyuan, Liu, Fayao, Yang, Xulei, Wei, Jiacheng, Qiu, Lemiao, Li, Hongsheng, Lin, Guosheng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918299143503872
author Fang, Naiyu
Zhou, Zheyuan
Liu, Fayao
Yang, Xulei
Wei, Jiacheng
Qiu, Lemiao
Li, Hongsheng
Lin, Guosheng
author_facet Fang, Naiyu
Zhou, Zheyuan
Liu, Fayao
Yang, Xulei
Wei, Jiacheng
Qiu, Lemiao
Li, Hongsheng
Lin, Guosheng
contents 3D semantic occupancy prediction offers an intuitive and efficient scene understanding and has attracted significant interest in autonomous driving perception. Existing approaches either rely on full supervision, which demands costly voxel-level annotations, or on self-supervision, which provides limited guidance and yields suboptimal performance. To address these challenges, we propose OccLE, a Label-Efficient 3D Semantic Occupancy Prediction that takes images and LiDAR as inputs and maintains high performance with limited voxel annotations. Our intuition is to decouple the semantic and geometric learning tasks and then fuse the learned feature grids from both tasks for the final semantic occupancy prediction. Therefore, the semantic branch distills 2D foundation model to provide aligned pseudo labels for 2D and 3D semantic learning. The geometric branch integrates image and LiDAR inputs in cross-plane synergy based on their inherency, employing semi-supervision to enhance geometry learning. We fuse semantic-geometric feature grids through Dual Mamba and incorporate a scatter-accumulated projection to supervise unannotated prediction with aligned pseudo labels. Experiments show that OccLE achieves competitive performance with only 10\% of voxel annotations on the SemanticKITTI and Occ3D-nuScenes datasets. The code will be publicly released on https://github.com/NerdFNY/OccLE
format Preprint
id arxiv_https___arxiv_org_abs_2505_20617
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle OccLE: Label-Efficient 3D Semantic Occupancy Prediction
Fang, Naiyu
Zhou, Zheyuan
Liu, Fayao
Yang, Xulei
Wei, Jiacheng
Qiu, Lemiao
Li, Hongsheng
Lin, Guosheng
Computer Vision and Pattern Recognition
3D semantic occupancy prediction offers an intuitive and efficient scene understanding and has attracted significant interest in autonomous driving perception. Existing approaches either rely on full supervision, which demands costly voxel-level annotations, or on self-supervision, which provides limited guidance and yields suboptimal performance. To address these challenges, we propose OccLE, a Label-Efficient 3D Semantic Occupancy Prediction that takes images and LiDAR as inputs and maintains high performance with limited voxel annotations. Our intuition is to decouple the semantic and geometric learning tasks and then fuse the learned feature grids from both tasks for the final semantic occupancy prediction. Therefore, the semantic branch distills 2D foundation model to provide aligned pseudo labels for 2D and 3D semantic learning. The geometric branch integrates image and LiDAR inputs in cross-plane synergy based on their inherency, employing semi-supervision to enhance geometry learning. We fuse semantic-geometric feature grids through Dual Mamba and incorporate a scatter-accumulated projection to supervise unannotated prediction with aligned pseudo labels. Experiments show that OccLE achieves competitive performance with only 10\% of voxel annotations on the SemanticKITTI and Occ3D-nuScenes datasets. The code will be publicly released on https://github.com/NerdFNY/OccLE
title OccLE: Label-Efficient 3D Semantic Occupancy Prediction
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.20617