SCube: Instant Large-Scale Scene Reconstruction using VoxSplats

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ren, Xuanchi, Lu, Yifan, Liang, Hanxue, Wu, Zhangjie, Ling, Huan, Chen, Mike, Fidler, Sanja, Williams, Francis, Huang, Jiahui
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916455853850624
author Ren, Xuanchi
Lu, Yifan
Liang, Hanxue
Wu, Zhangjie
Ling, Huan
Chen, Mike
Fidler, Sanja
Williams, Francis
Huang, Jiahui
author_facet Ren, Xuanchi
Lu, Yifan
Liang, Hanxue
Wu, Zhangjie
Ling, Huan
Chen, Mike
Fidler, Sanja
Williams, Francis
Huang, Jiahui
contents We present SCube, a novel method for reconstructing large-scale 3D scenes (geometry, appearance, and semantics) from a sparse set of posed images. Our method encodes reconstructed scenes using a novel representation VoxSplat, which is a set of 3D Gaussians supported on a high-resolution sparse-voxel scaffold. To reconstruct a VoxSplat from images, we employ a hierarchical voxel latent diffusion model conditioned on the input images followed by a feedforward appearance prediction model. The diffusion model generates high-resolution grids progressively in a coarse-to-fine manner, and the appearance network predicts a set of Gaussians within each voxel. From as few as 3 non-overlapping input images, SCube can generate millions of Gaussians with a 1024^3 voxel grid spanning hundreds of meters in 20 seconds. Past works tackling scene reconstruction from images either rely on per-scene optimization and fail to reconstruct the scene away from input views (thus requiring dense view coverage as input) or leverage geometric priors based on low-resolution models, which produce blurry results. In contrast, SCube leverages high-resolution sparse networks and produces sharp outputs from few views. We show the superiority of SCube compared to prior art using the Waymo self-driving dataset on 3D reconstruction and demonstrate its applications, such as LiDAR simulation and text-to-scene generation.
format Preprint
id arxiv_https___arxiv_org_abs_2410_20030
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle SCube: Instant Large-Scale Scene Reconstruction using VoxSplats
Ren, Xuanchi
Lu, Yifan
Liang, Hanxue
Wu, Zhangjie
Ling, Huan
Chen, Mike
Fidler, Sanja
Williams, Francis
Huang, Jiahui
Computer Vision and Pattern Recognition
Artificial Intelligence
Graphics
We present SCube, a novel method for reconstructing large-scale 3D scenes (geometry, appearance, and semantics) from a sparse set of posed images. Our method encodes reconstructed scenes using a novel representation VoxSplat, which is a set of 3D Gaussians supported on a high-resolution sparse-voxel scaffold. To reconstruct a VoxSplat from images, we employ a hierarchical voxel latent diffusion model conditioned on the input images followed by a feedforward appearance prediction model. The diffusion model generates high-resolution grids progressively in a coarse-to-fine manner, and the appearance network predicts a set of Gaussians within each voxel. From as few as 3 non-overlapping input images, SCube can generate millions of Gaussians with a 1024^3 voxel grid spanning hundreds of meters in 20 seconds. Past works tackling scene reconstruction from images either rely on per-scene optimization and fail to reconstruct the scene away from input views (thus requiring dense view coverage as input) or leverage geometric priors based on low-resolution models, which produce blurry results. In contrast, SCube leverages high-resolution sparse networks and produces sharp outputs from few views. We show the superiority of SCube compared to prior art using the Waymo self-driving dataset on 3D reconstruction and demonstrate its applications, such as LiDAR simulation and text-to-scene generation.
title SCube: Instant Large-Scale Scene Reconstruction using VoxSplats
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Graphics
url https://arxiv.org/abs/2410.20030