XCube: Large-Scale 3D Generative Modeling using Sparse Voxel Hierarchies

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ren, Xuanchi, Huang, Jiahui, Zeng, Xiaohui, Museth, Ken, Fidler, Sanja, Williams, Francis
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916299535286272
author Ren, Xuanchi
Huang, Jiahui
Zeng, Xiaohui
Museth, Ken
Fidler, Sanja
Williams, Francis
author_facet Ren, Xuanchi
Huang, Jiahui
Zeng, Xiaohui
Museth, Ken
Fidler, Sanja
Williams, Francis
contents We present XCube (abbreviated as $\mathcal{X}^3$), a novel generative model for high-resolution sparse 3D voxel grids with arbitrary attributes. Our model can generate millions of voxels with a finest effective resolution of up to $1024^3$ in a feed-forward fashion without time-consuming test-time optimization. To achieve this, we employ a hierarchical voxel latent diffusion model which generates progressively higher resolution grids in a coarse-to-fine manner using a custom framework built on the highly efficient VDB data structure. Apart from generating high-resolution objects, we demonstrate the effectiveness of XCube on large outdoor scenes at scales of 100m$\times$100m with a voxel size as small as 10cm. We observe clear qualitative and quantitative improvements over past approaches. In addition to unconditional generation, we show that our model can be used to solve a variety of tasks such as user-guided editing, scene completion from a single scan, and text-to-3D. The source code and more results can be found at https://research.nvidia.com/labs/toronto-ai/xcube/.
format Preprint
id arxiv_https___arxiv_org_abs_2312_03806
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle XCube: Large-Scale 3D Generative Modeling using Sparse Voxel Hierarchies
Ren, Xuanchi
Huang, Jiahui
Zeng, Xiaohui
Museth, Ken
Fidler, Sanja
Williams, Francis
Computer Vision and Pattern Recognition
Graphics
Machine Learning
We present XCube (abbreviated as $\mathcal{X}^3$), a novel generative model for high-resolution sparse 3D voxel grids with arbitrary attributes. Our model can generate millions of voxels with a finest effective resolution of up to $1024^3$ in a feed-forward fashion without time-consuming test-time optimization. To achieve this, we employ a hierarchical voxel latent diffusion model which generates progressively higher resolution grids in a coarse-to-fine manner using a custom framework built on the highly efficient VDB data structure. Apart from generating high-resolution objects, we demonstrate the effectiveness of XCube on large outdoor scenes at scales of 100m$\times$100m with a voxel size as small as 10cm. We observe clear qualitative and quantitative improvements over past approaches. In addition to unconditional generation, we show that our model can be used to solve a variety of tasks such as user-guided editing, scene completion from a single scan, and text-to-3D. The source code and more results can be found at https://research.nvidia.com/labs/toronto-ai/xcube/.
title XCube: Large-Scale 3D Generative Modeling using Sparse Voxel Hierarchies
topic Computer Vision and Pattern Recognition
Graphics
Machine Learning
url https://arxiv.org/abs/2312.03806