DreamTime: An Improved Optimization Strategy for Diffusion-Guided 3D Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, Yukun, Wang, Jianan, Shi, Yukai, Tang, Boshi, Qi, Xianbiao, Zhang, Lei
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917657607929856
author Huang, Yukun
Wang, Jianan
Shi, Yukai
Tang, Boshi
Qi, Xianbiao
Zhang, Lei
author_facet Huang, Yukun
Wang, Jianan
Shi, Yukai
Tang, Boshi
Qi, Xianbiao
Zhang, Lei
contents Text-to-image diffusion models pre-trained on billions of image-text pairs have recently enabled 3D content creation by optimizing a randomly initialized differentiable 3D representation with score distillation. However, the optimization process suffers slow convergence and the resultant 3D models often exhibit two limitations: (a) quality concerns such as missing attributes and distorted shape and texture; (b) extremely low diversity comparing to text-guided image synthesis. In this paper, we show that the conflict between the 3D optimization process and uniform timestep sampling in score distillation is the main reason for these limitations. To resolve this conflict, we propose to prioritize timestep sampling with monotonically non-increasing functions, which aligns the 3D optimization process with the sampling process of diffusion model. Extensive experiments show that our simple redesign significantly improves 3D content creation with faster convergence, better quality and diversity.
format Preprint
id arxiv_https___arxiv_org_abs_2306_12422
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle DreamTime: An Improved Optimization Strategy for Diffusion-Guided 3D Generation
Huang, Yukun
Wang, Jianan
Shi, Yukai
Tang, Boshi
Qi, Xianbiao
Zhang, Lei
Computer Vision and Pattern Recognition
Graphics
Machine Learning
Text-to-image diffusion models pre-trained on billions of image-text pairs have recently enabled 3D content creation by optimizing a randomly initialized differentiable 3D representation with score distillation. However, the optimization process suffers slow convergence and the resultant 3D models often exhibit two limitations: (a) quality concerns such as missing attributes and distorted shape and texture; (b) extremely low diversity comparing to text-guided image synthesis. In this paper, we show that the conflict between the 3D optimization process and uniform timestep sampling in score distillation is the main reason for these limitations. To resolve this conflict, we propose to prioritize timestep sampling with monotonically non-increasing functions, which aligns the 3D optimization process with the sampling process of diffusion model. Extensive experiments show that our simple redesign significantly improves 3D content creation with faster convergence, better quality and diversity.
title DreamTime: An Improved Optimization Strategy for Diffusion-Guided 3D Generation
topic Computer Vision and Pattern Recognition
Graphics
Machine Learning
url https://arxiv.org/abs/2306.12422