GS-Occ3D: Scaling Vision-only Occupancy Reconstruction with Gaussian Splatting

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ye, Baijun, Qin, Minghui, Zhang, Saining, Gong, Moonjun, Zhu, Shaoting, Shen, Zebang, Zhang, Luan, Zhang, Lu, Zhao, Hao, Zhao, Hang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908476509257728
author Ye, Baijun
Qin, Minghui
Zhang, Saining
Gong, Moonjun
Zhu, Shaoting
Shen, Zebang
Zhang, Luan
Zhang, Lu
Zhao, Hao
Zhao, Hang
author_facet Ye, Baijun
Qin, Minghui
Zhang, Saining
Gong, Moonjun
Zhu, Shaoting
Shen, Zebang
Zhang, Luan
Zhang, Lu
Zhao, Hao
Zhao, Hang
contents Occupancy is crucial for autonomous driving, providing essential geometric priors for perception and planning. However, existing methods predominantly rely on LiDAR-based occupancy annotations, which limits scalability and prevents leveraging vast amounts of potential crowdsourced data for auto-labeling. To address this, we propose GS-Occ3D, a scalable vision-only framework that directly reconstructs occupancy. Vision-only occupancy reconstruction poses significant challenges due to sparse viewpoints, dynamic scene elements, severe occlusions, and long-horizon motion. Existing vision-based methods primarily rely on mesh representation, which suffer from incomplete geometry and additional post-processing, limiting scalability. To overcome these issues, GS-Occ3D optimizes an explicit occupancy representation using an Octree-based Gaussian Surfel formulation, ensuring efficiency and scalability. Additionally, we decompose scenes into static background, ground, and dynamic objects, enabling tailored modeling strategies: (1) Ground is explicitly reconstructed as a dominant structural element, significantly improving large-area consistency; (2) Dynamic vehicles are separately modeled to better capture motion-related occupancy patterns. Extensive experiments on the Waymo dataset demonstrate that GS-Occ3D achieves state-of-the-art geometry reconstruction results. By curating vision-only binary occupancy labels from diverse urban scenes, we show their effectiveness for downstream occupancy models on Occ3D-Waymo and superior zero-shot generalization on Occ3D-nuScenes. It highlights the potential of large-scale vision-based occupancy reconstruction as a new paradigm for scalable auto-labeling. Project Page: https://gs-occ3d.github.io/
format Preprint
id arxiv_https___arxiv_org_abs_2507_19451
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle GS-Occ3D: Scaling Vision-only Occupancy Reconstruction with Gaussian Splatting
Ye, Baijun
Qin, Minghui
Zhang, Saining
Gong, Moonjun
Zhu, Shaoting
Shen, Zebang
Zhang, Luan
Zhang, Lu
Zhao, Hao
Zhao, Hang
Computer Vision and Pattern Recognition
Occupancy is crucial for autonomous driving, providing essential geometric priors for perception and planning. However, existing methods predominantly rely on LiDAR-based occupancy annotations, which limits scalability and prevents leveraging vast amounts of potential crowdsourced data for auto-labeling. To address this, we propose GS-Occ3D, a scalable vision-only framework that directly reconstructs occupancy. Vision-only occupancy reconstruction poses significant challenges due to sparse viewpoints, dynamic scene elements, severe occlusions, and long-horizon motion. Existing vision-based methods primarily rely on mesh representation, which suffer from incomplete geometry and additional post-processing, limiting scalability. To overcome these issues, GS-Occ3D optimizes an explicit occupancy representation using an Octree-based Gaussian Surfel formulation, ensuring efficiency and scalability. Additionally, we decompose scenes into static background, ground, and dynamic objects, enabling tailored modeling strategies: (1) Ground is explicitly reconstructed as a dominant structural element, significantly improving large-area consistency; (2) Dynamic vehicles are separately modeled to better capture motion-related occupancy patterns. Extensive experiments on the Waymo dataset demonstrate that GS-Occ3D achieves state-of-the-art geometry reconstruction results. By curating vision-only binary occupancy labels from diverse urban scenes, we show their effectiveness for downstream occupancy models on Occ3D-Waymo and superior zero-shot generalization on Occ3D-nuScenes. It highlights the potential of large-scale vision-based occupancy reconstruction as a new paradigm for scalable auto-labeling. Project Page: https://gs-occ3d.github.io/
title GS-Occ3D: Scaling Vision-only Occupancy Reconstruction with Gaussian Splatting
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2507.19451