SCube: Instant Large-Scale Scene Reconstruction using VoxSplats
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866916455853850624 |
|---|---|
| author | Ren, Xuanchi Lu, Yifan Liang, Hanxue Wu, Zhangjie Ling, Huan Chen, Mike Fidler, Sanja Williams, Francis Huang, Jiahui |
| author_facet | Ren, Xuanchi Lu, Yifan Liang, Hanxue Wu, Zhangjie Ling, Huan Chen, Mike Fidler, Sanja Williams, Francis Huang, Jiahui |
| contents | We present SCube, a novel method for reconstructing large-scale 3D scenes (geometry, appearance, and semantics) from a sparse set of posed images. Our method encodes reconstructed scenes using a novel representation VoxSplat, which is a set of 3D Gaussians supported on a high-resolution sparse-voxel scaffold. To reconstruct a VoxSplat from images, we employ a hierarchical voxel latent diffusion model conditioned on the input images followed by a feedforward appearance prediction model. The diffusion model generates high-resolution grids progressively in a coarse-to-fine manner, and the appearance network predicts a set of Gaussians within each voxel. From as few as 3 non-overlapping input images, SCube can generate millions of Gaussians with a 1024^3 voxel grid spanning hundreds of meters in 20 seconds. Past works tackling scene reconstruction from images either rely on per-scene optimization and fail to reconstruct the scene away from input views (thus requiring dense view coverage as input) or leverage geometric priors based on low-resolution models, which produce blurry results. In contrast, SCube leverages high-resolution sparse networks and produces sharp outputs from few views. We show the superiority of SCube compared to prior art using the Waymo self-driving dataset on 3D reconstruction and demonstrate its applications, such as LiDAR simulation and text-to-scene generation. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2410_20030 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | SCube: Instant Large-Scale Scene Reconstruction using VoxSplats Ren, Xuanchi Lu, Yifan Liang, Hanxue Wu, Zhangjie Ling, Huan Chen, Mike Fidler, Sanja Williams, Francis Huang, Jiahui Computer Vision and Pattern Recognition Artificial Intelligence Graphics We present SCube, a novel method for reconstructing large-scale 3D scenes (geometry, appearance, and semantics) from a sparse set of posed images. Our method encodes reconstructed scenes using a novel representation VoxSplat, which is a set of 3D Gaussians supported on a high-resolution sparse-voxel scaffold. To reconstruct a VoxSplat from images, we employ a hierarchical voxel latent diffusion model conditioned on the input images followed by a feedforward appearance prediction model. The diffusion model generates high-resolution grids progressively in a coarse-to-fine manner, and the appearance network predicts a set of Gaussians within each voxel. From as few as 3 non-overlapping input images, SCube can generate millions of Gaussians with a 1024^3 voxel grid spanning hundreds of meters in 20 seconds. Past works tackling scene reconstruction from images either rely on per-scene optimization and fail to reconstruct the scene away from input views (thus requiring dense view coverage as input) or leverage geometric priors based on low-resolution models, which produce blurry results. In contrast, SCube leverages high-resolution sparse networks and produces sharp outputs from few views. We show the superiority of SCube compared to prior art using the Waymo self-driving dataset on 3D reconstruction and demonstrate its applications, such as LiDAR simulation and text-to-scene generation. |
| title | SCube: Instant Large-Scale Scene Reconstruction using VoxSplats |
| topic | Computer Vision and Pattern Recognition Artificial Intelligence Graphics |
| url | https://arxiv.org/abs/2410.20030 |