ALOcc: Adaptive Lifting-Based 3D Semantic Occupancy and Cost Volume-Based Flow Predictions

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Dubing, Fang, Jin, Han, Wencheng, Cheng, Xinjing, Yin, Junbo, Xu, Chenzhong, Khan, Fahad Shahbaz, Shen, Jianbing
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908529164550144
author Chen, Dubing
Fang, Jin
Han, Wencheng
Cheng, Xinjing
Yin, Junbo
Xu, Chenzhong
Khan, Fahad Shahbaz
Shen, Jianbing
author_facet Chen, Dubing
Fang, Jin
Han, Wencheng
Cheng, Xinjing
Yin, Junbo
Xu, Chenzhong
Khan, Fahad Shahbaz
Shen, Jianbing
contents 3D semantic occupancy and flow prediction are fundamental to spatiotemporal scene understanding. This paper proposes a vision-based framework with three targeted improvements. First, we introduce an occlusion-aware adaptive lifting mechanism incorporating depth denoising. This enhances the robustness of 2D-to-3D feature transformation while mitigating reliance on depth priors. Second, we enforce 3D-2D semantic consistency via jointly optimized prototypes, using confidence- and category-aware sampling to address the long-tail classes problem. Third, to streamline joint prediction, we devise a BEV-centric cost volume to explicitly correlate semantic and flow features, supervised by a hybrid classification-regression scheme that handles diverse motion scales. Our purely convolutional architecture establishes new SOTA performance on multiple benchmarks for both semantic occupancy and joint occupancy semantic-flow prediction. We also present a family of models offering a spectrum of efficiency-performance trade-offs. Our real-time version exceeds all existing real-time methods in speed and accuracy, ensuring its practical viability.
format Preprint
id arxiv_https___arxiv_org_abs_2411_07725
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle ALOcc: Adaptive Lifting-Based 3D Semantic Occupancy and Cost Volume-Based Flow Predictions
Chen, Dubing
Fang, Jin
Han, Wencheng
Cheng, Xinjing
Yin, Junbo
Xu, Chenzhong
Khan, Fahad Shahbaz
Shen, Jianbing
Computer Vision and Pattern Recognition
3D semantic occupancy and flow prediction are fundamental to spatiotemporal scene understanding. This paper proposes a vision-based framework with three targeted improvements. First, we introduce an occlusion-aware adaptive lifting mechanism incorporating depth denoising. This enhances the robustness of 2D-to-3D feature transformation while mitigating reliance on depth priors. Second, we enforce 3D-2D semantic consistency via jointly optimized prototypes, using confidence- and category-aware sampling to address the long-tail classes problem. Third, to streamline joint prediction, we devise a BEV-centric cost volume to explicitly correlate semantic and flow features, supervised by a hybrid classification-regression scheme that handles diverse motion scales. Our purely convolutional architecture establishes new SOTA performance on multiple benchmarks for both semantic occupancy and joint occupancy semantic-flow prediction. We also present a family of models offering a spectrum of efficiency-performance trade-offs. Our real-time version exceeds all existing real-time methods in speed and accuracy, ensuring its practical viability.
title ALOcc: Adaptive Lifting-Based 3D Semantic Occupancy and Cost Volume-Based Flow Predictions
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2411.07725