Progressive3D: Progressively Local Editing for Text-to-3D Content Creation with Complex Semantic Prompts

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cheng, Xinhua, Yang, Tianyu, Wang, Jianan, Li, Yu, Zhang, Lei, Zhang, Jian, Yuan, Li
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913266989531136
author Cheng, Xinhua
Yang, Tianyu
Wang, Jianan
Li, Yu
Zhang, Lei
Zhang, Jian
Yuan, Li
author_facet Cheng, Xinhua
Yang, Tianyu
Wang, Jianan
Li, Yu
Zhang, Lei
Zhang, Jian
Yuan, Li
contents Recent text-to-3D generation methods achieve impressive 3D content creation capacity thanks to the advances in image diffusion models and optimizing strategies. However, current methods struggle to generate correct 3D content for a complex prompt in semantics, i.e., a prompt describing multiple interacted objects binding with different attributes. In this work, we propose a general framework named Progressive3D, which decomposes the entire generation into a series of locally progressive editing steps to create precise 3D content for complex prompts, and we constrain the content change to only occur in regions determined by user-defined region prompts in each editing step. Furthermore, we propose an overlapped semantic component suppression technique to encourage the optimization process to focus more on the semantic differences between prompts. Extensive experiments demonstrate that the proposed Progressive3D framework generates precise 3D content for prompts with complex semantics and is general for various text-to-3D methods driven by different 3D representations.
format Preprint
id arxiv_https___arxiv_org_abs_2310_11784
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Progressive3D: Progressively Local Editing for Text-to-3D Content Creation with Complex Semantic Prompts
Cheng, Xinhua
Yang, Tianyu
Wang, Jianan
Li, Yu
Zhang, Lei
Zhang, Jian
Yuan, Li
Computer Vision and Pattern Recognition
Recent text-to-3D generation methods achieve impressive 3D content creation capacity thanks to the advances in image diffusion models and optimizing strategies. However, current methods struggle to generate correct 3D content for a complex prompt in semantics, i.e., a prompt describing multiple interacted objects binding with different attributes. In this work, we propose a general framework named Progressive3D, which decomposes the entire generation into a series of locally progressive editing steps to create precise 3D content for complex prompts, and we constrain the content change to only occur in regions determined by user-defined region prompts in each editing step. Furthermore, we propose an overlapped semantic component suppression technique to encourage the optimization process to focus more on the semantic differences between prompts. Extensive experiments demonstrate that the proposed Progressive3D framework generates precise 3D content for prompts with complex semantics and is general for various text-to-3D methods driven by different 3D representations.
title Progressive3D: Progressively Local Editing for Text-to-3D Content Creation with Complex Semantic Prompts
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2310.11784