VolumeDiffusion: Flexible Text-to-3D Generation with Efficient Volumetric Encoder

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tang, Zhicong, Gu, Shuyang, Wang, Chunyu, Zhang, Ting, Bao, Jianmin, Chen, Dong, Guo, Baining
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917747204554752
author Tang, Zhicong
Gu, Shuyang
Wang, Chunyu
Zhang, Ting
Bao, Jianmin
Chen, Dong
Guo, Baining
author_facet Tang, Zhicong
Gu, Shuyang
Wang, Chunyu
Zhang, Ting
Bao, Jianmin
Chen, Dong
Guo, Baining
contents This paper introduces a pioneering 3D volumetric encoder designed for text-to-3D generation. To scale up the training data for the diffusion model, a lightweight network is developed to efficiently acquire feature volumes from multi-view images. The 3D volumes are then trained on a diffusion model for text-to-3D generation using a 3D U-Net. This research further addresses the challenges of inaccurate object captions and high-dimensional feature volumes. The proposed model, trained on the public Objaverse dataset, demonstrates promising outcomes in producing diverse and recognizable samples from text prompts. Notably, it empowers finer control over object part characteristics through textual cues, fostering model creativity by seamlessly combining multiple concepts within a single object. This research significantly contributes to the progress of 3D generation by introducing an efficient, flexible, and scalable representation methodology.
format Preprint
id arxiv_https___arxiv_org_abs_2312_11459
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle VolumeDiffusion: Flexible Text-to-3D Generation with Efficient Volumetric Encoder
Tang, Zhicong
Gu, Shuyang
Wang, Chunyu
Zhang, Ting
Bao, Jianmin
Chen, Dong
Guo, Baining
Computer Vision and Pattern Recognition
This paper introduces a pioneering 3D volumetric encoder designed for text-to-3D generation. To scale up the training data for the diffusion model, a lightweight network is developed to efficiently acquire feature volumes from multi-view images. The 3D volumes are then trained on a diffusion model for text-to-3D generation using a 3D U-Net. This research further addresses the challenges of inaccurate object captions and high-dimensional feature volumes. The proposed model, trained on the public Objaverse dataset, demonstrates promising outcomes in producing diverse and recognizable samples from text prompts. Notably, it empowers finer control over object part characteristics through textual cues, fostering model creativity by seamlessly combining multiple concepts within a single object. This research significantly contributes to the progress of 3D generation by introducing an efficient, flexible, and scalable representation methodology.
title VolumeDiffusion: Flexible Text-to-3D Generation with Efficient Volumetric Encoder
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2312.11459