ExactDreamer: High-Fidelity Text-to-3D Content Creation via Exact Score Matching

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Yumin, Miao, Xingyu, Duan, Haoran, Wei, Bo, Shah, Tejal, Long, Yang, Ranjan, Rajiv
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916259675766784
author Zhang, Yumin
Miao, Xingyu
Duan, Haoran
Wei, Bo
Shah, Tejal
Long, Yang
Ranjan, Rajiv
author_facet Zhang, Yumin
Miao, Xingyu
Duan, Haoran
Wei, Bo
Shah, Tejal
Long, Yang
Ranjan, Rajiv
contents Text-to-3D content creation is a rapidly evolving research area. Given the scarcity of 3D data, current approaches often adapt pre-trained 2D diffusion models for 3D synthesis. Among these approaches, Score Distillation Sampling (SDS) has been widely adopted. However, the issue of over-smoothing poses a significant limitation on the high-fidelity generation of 3D models. To address this challenge, LucidDreamer replaces the Denoising Diffusion Probabilistic Model (DDPM) in SDS with the Denoising Diffusion Implicit Model (DDIM) to construct Interval Score Matching (ISM). However, ISM inevitably inherits inconsistencies from DDIM, causing reconstruction errors during the DDIM inversion process. This results in poor performance in the detailed generation of 3D objects and loss of content. To alleviate these problems, we propose a novel method named Exact Score Matching (ESM). Specifically, ESM leverages auxiliary variables to mathematically guarantee exact recovery in the DDIM reverse process. Furthermore, to effectively capture the dynamic changes of the original and auxiliary variables, the LoRA of a pre-trained diffusion model implements these exact paths. Extensive experiments demonstrate the effectiveness of ESM in text-to-3D generation, particularly highlighting its superiority in detailed generation.
format Preprint
id arxiv_https___arxiv_org_abs_2405_15914
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle ExactDreamer: High-Fidelity Text-to-3D Content Creation via Exact Score Matching
Zhang, Yumin
Miao, Xingyu
Duan, Haoran
Wei, Bo
Shah, Tejal
Long, Yang
Ranjan, Rajiv
Computer Vision and Pattern Recognition
Text-to-3D content creation is a rapidly evolving research area. Given the scarcity of 3D data, current approaches often adapt pre-trained 2D diffusion models for 3D synthesis. Among these approaches, Score Distillation Sampling (SDS) has been widely adopted. However, the issue of over-smoothing poses a significant limitation on the high-fidelity generation of 3D models. To address this challenge, LucidDreamer replaces the Denoising Diffusion Probabilistic Model (DDPM) in SDS with the Denoising Diffusion Implicit Model (DDIM) to construct Interval Score Matching (ISM). However, ISM inevitably inherits inconsistencies from DDIM, causing reconstruction errors during the DDIM inversion process. This results in poor performance in the detailed generation of 3D objects and loss of content. To alleviate these problems, we propose a novel method named Exact Score Matching (ESM). Specifically, ESM leverages auxiliary variables to mathematically guarantee exact recovery in the DDIM reverse process. Furthermore, to effectively capture the dynamic changes of the original and auxiliary variables, the LoRA of a pre-trained diffusion model implements these exact paths. Extensive experiments demonstrate the effectiveness of ESM in text-to-3D generation, particularly highlighting its superiority in detailed generation.
title ExactDreamer: High-Fidelity Text-to-3D Content Creation via Exact Score Matching
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2405.15914