Direct2.5: Diverse Text-to-3D Generation via Multi-view 2.5D Diffusion

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lu, Yuanxun, Zhang, Jingyang, Li, Shiwei, Fang, Tian, McKinnon, David, Tsin, Yanghai, Quan, Long, Cao, Xun, Yao, Yao
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909144188977152
author Lu, Yuanxun
Zhang, Jingyang
Li, Shiwei
Fang, Tian
McKinnon, David
Tsin, Yanghai
Quan, Long
Cao, Xun
Yao, Yao
author_facet Lu, Yuanxun
Zhang, Jingyang
Li, Shiwei
Fang, Tian
McKinnon, David
Tsin, Yanghai
Quan, Long
Cao, Xun
Yao, Yao
contents Recent advances in generative AI have unveiled significant potential for the creation of 3D content. However, current methods either apply a pre-trained 2D diffusion model with the time-consuming score distillation sampling (SDS), or a direct 3D diffusion model trained on limited 3D data losing generation diversity. In this work, we approach the problem by employing a multi-view 2.5D diffusion fine-tuned from a pre-trained 2D diffusion model. The multi-view 2.5D diffusion directly models the structural distribution of 3D data, while still maintaining the strong generalization ability of the original 2D diffusion model, filling the gap between 2D diffusion-based and direct 3D diffusion-based methods for 3D content generation. During inference, multi-view normal maps are generated using the 2.5D diffusion, and a novel differentiable rasterization scheme is introduced to fuse the almost consistent multi-view normal maps into a consistent 3D model. We further design a normal-conditioned multi-view image generation module for fast appearance generation given the 3D geometry. Our method is a one-pass diffusion process and does not require any SDS optimization as post-processing. We demonstrate through extensive experiments that, our direct 2.5D generation with the specially-designed fusion scheme can achieve diverse, mode-seeking-free, and high-fidelity 3D content generation in only 10 seconds. Project page: https://nju-3dv.github.io/projects/direct25.
format Preprint
id arxiv_https___arxiv_org_abs_2311_15980
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Direct2.5: Diverse Text-to-3D Generation via Multi-view 2.5D Diffusion
Lu, Yuanxun
Zhang, Jingyang
Li, Shiwei
Fang, Tian
McKinnon, David
Tsin, Yanghai
Quan, Long
Cao, Xun
Yao, Yao
Computer Vision and Pattern Recognition
Recent advances in generative AI have unveiled significant potential for the creation of 3D content. However, current methods either apply a pre-trained 2D diffusion model with the time-consuming score distillation sampling (SDS), or a direct 3D diffusion model trained on limited 3D data losing generation diversity. In this work, we approach the problem by employing a multi-view 2.5D diffusion fine-tuned from a pre-trained 2D diffusion model. The multi-view 2.5D diffusion directly models the structural distribution of 3D data, while still maintaining the strong generalization ability of the original 2D diffusion model, filling the gap between 2D diffusion-based and direct 3D diffusion-based methods for 3D content generation. During inference, multi-view normal maps are generated using the 2.5D diffusion, and a novel differentiable rasterization scheme is introduced to fuse the almost consistent multi-view normal maps into a consistent 3D model. We further design a normal-conditioned multi-view image generation module for fast appearance generation given the 3D geometry. Our method is a one-pass diffusion process and does not require any SDS optimization as post-processing. We demonstrate through extensive experiments that, our direct 2.5D generation with the specially-designed fusion scheme can achieve diverse, mode-seeking-free, and high-fidelity 3D content generation in only 10 seconds. Project page: https://nju-3dv.github.io/projects/direct25.
title Direct2.5: Diverse Text-to-3D Generation via Multi-view 2.5D Diffusion
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2311.15980