Stream and Query-guided Feature Aggregation for Efficient and Effective 3D Occupancy Prediction

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Moon, Seokha, Baek, Janghyun, Kim, Giseop, Kim, Jinkyu, Choi, Sunwook
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909926019825664
author Moon, Seokha
Baek, Janghyun
Kim, Giseop
Kim, Jinkyu
Choi, Sunwook
author_facet Moon, Seokha
Baek, Janghyun
Kim, Giseop
Kim, Jinkyu
Choi, Sunwook
contents 3D occupancy prediction has become a key perception task in autonomous driving, as it enables comprehensive scene understanding. Recent methods enhance this understanding by incorporating spatiotemporal information through multi-frame fusion, but they suffer from a trade-off: dense voxel-based representations provide high accuracy at significant computational cost, whereas sparse representations improve efficiency but lose spatial detail. To mitigate this trade-off, we introduce DuOcc, which employs a dual aggregation strategy that retains dense voxel representations to preserve spatial fidelity while maintaining high efficiency. DuOcc consists of two key components: (i) Stream-based Voxel Aggregation, which recurrently accumulates voxel features over time and refines them to suppress warping-induced distortions, preserving a clear separation between occupied and free space. (ii) Query-guided Aggregation, which complements the limitations of voxel accumulation by selectively injecting instance-level query features into the voxel regions occupied by dynamic objects. Experiments on the widely used Occ3D-nuScenes and SurroundOcc datasets demonstrate that DuOcc achieves state-of-the-art performance in real-time settings, while reducing memory usage by over 40% compared to prior methods.
format Preprint
id arxiv_https___arxiv_org_abs_2503_22087
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Stream and Query-guided Feature Aggregation for Efficient and Effective 3D Occupancy Prediction
Moon, Seokha
Baek, Janghyun
Kim, Giseop
Kim, Jinkyu
Choi, Sunwook
Computer Vision and Pattern Recognition
3D occupancy prediction has become a key perception task in autonomous driving, as it enables comprehensive scene understanding. Recent methods enhance this understanding by incorporating spatiotemporal information through multi-frame fusion, but they suffer from a trade-off: dense voxel-based representations provide high accuracy at significant computational cost, whereas sparse representations improve efficiency but lose spatial detail. To mitigate this trade-off, we introduce DuOcc, which employs a dual aggregation strategy that retains dense voxel representations to preserve spatial fidelity while maintaining high efficiency. DuOcc consists of two key components: (i) Stream-based Voxel Aggregation, which recurrently accumulates voxel features over time and refines them to suppress warping-induced distortions, preserving a clear separation between occupied and free space. (ii) Query-guided Aggregation, which complements the limitations of voxel accumulation by selectively injecting instance-level query features into the voxel regions occupied by dynamic objects. Experiments on the widely used Occ3D-nuScenes and SurroundOcc datasets demonstrate that DuOcc achieves state-of-the-art performance in real-time settings, while reducing memory usage by over 40% compared to prior methods.
title Stream and Query-guided Feature Aggregation for Efficient and Effective 3D Occupancy Prediction
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.22087