QuadricFormer: Scene as Superquadrics for 3D Semantic Occupancy Prediction

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zuo, Sicheng, Zheng, Wenzhao, Han, Xiaoyong, Yang, Longchao, Pan, Yong, Lu, Jiwen
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916792194039808
author Zuo, Sicheng
Zheng, Wenzhao
Han, Xiaoyong
Yang, Longchao
Pan, Yong
Lu, Jiwen
author_facet Zuo, Sicheng
Zheng, Wenzhao
Han, Xiaoyong
Yang, Longchao
Pan, Yong
Lu, Jiwen
contents 3D occupancy prediction is crucial for robust autonomous driving systems as it enables comprehensive perception of environmental structures and semantics. Most existing methods employ dense voxel-based scene representations, ignoring the sparsity of driving scenes and resulting in inefficiency. Recent works explore object-centric representations based on sparse Gaussians, but their ellipsoidal shape prior limits the modeling of diverse structures. In real-world driving scenes, objects exhibit rich geometries (e.g., cuboids, cylinders, and irregular shapes), necessitating excessive ellipsoidal Gaussians densely packed for accurate modeling, which leads to inefficient representations. To address this, we propose to use geometrically expressive superquadrics as scene primitives, enabling efficient representation of complex structures with fewer primitives through their inherent shape diversity. We develop a probabilistic superquadric mixture model, which interprets each superquadric as an occupancy probability distribution with a corresponding geometry prior, and calculates semantics through probabilistic mixture. Building on this, we present QuadricFormer, a superquadric-based model for efficient 3D occupancy prediction, and introduce a pruning-and-splitting module to further enhance modeling efficiency by concentrating superquadrics in occupied regions. Extensive experiments on the nuScenes dataset demonstrate that QuadricFormer achieves state-of-the-art performance while maintaining superior efficiency.
format Preprint
id arxiv_https___arxiv_org_abs_2506_10977
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle QuadricFormer: Scene as Superquadrics for 3D Semantic Occupancy Prediction
Zuo, Sicheng
Zheng, Wenzhao
Han, Xiaoyong
Yang, Longchao
Pan, Yong
Lu, Jiwen
Computer Vision and Pattern Recognition
3D occupancy prediction is crucial for robust autonomous driving systems as it enables comprehensive perception of environmental structures and semantics. Most existing methods employ dense voxel-based scene representations, ignoring the sparsity of driving scenes and resulting in inefficiency. Recent works explore object-centric representations based on sparse Gaussians, but their ellipsoidal shape prior limits the modeling of diverse structures. In real-world driving scenes, objects exhibit rich geometries (e.g., cuboids, cylinders, and irregular shapes), necessitating excessive ellipsoidal Gaussians densely packed for accurate modeling, which leads to inefficient representations. To address this, we propose to use geometrically expressive superquadrics as scene primitives, enabling efficient representation of complex structures with fewer primitives through their inherent shape diversity. We develop a probabilistic superquadric mixture model, which interprets each superquadric as an occupancy probability distribution with a corresponding geometry prior, and calculates semantics through probabilistic mixture. Building on this, we present QuadricFormer, a superquadric-based model for efficient 3D occupancy prediction, and introduce a pruning-and-splitting module to further enhance modeling efficiency by concentrating superquadrics in occupied regions. Extensive experiments on the nuScenes dataset demonstrate that QuadricFormer achieves state-of-the-art performance while maintaining superior efficiency.
title QuadricFormer: Scene as Superquadrics for 3D Semantic Occupancy Prediction
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.10977