TV-3DG: Mastering Text-to-3D Customized Generation with Visual Prompt

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Jiahui, Di, Donglin, Ma, Baorui, Yang, Xun, Ma, Yongjia, Sun, Wenzhang, Chen, Wei, Cui, Jianxun, Xue, Zhou, Wang, Meng, Liu, Yebin
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914999264346112
author Yang, Jiahui
Di, Donglin
Ma, Baorui
Yang, Xun
Ma, Yongjia
Sun, Wenzhang
Chen, Wei
Cui, Jianxun
Xue, Zhou
Wang, Meng
Liu, Yebin
author_facet Yang, Jiahui
Di, Donglin
Ma, Baorui
Yang, Xun
Ma, Yongjia
Sun, Wenzhang
Chen, Wei
Cui, Jianxun
Xue, Zhou
Wang, Meng
Liu, Yebin
contents In recent years, advancements in generative models have significantly expanded the capabilities of text-to-3D generation. Many approaches rely on Score Distillation Sampling (SDS) technology. However, SDS struggles to accommodate multi-condition inputs, such as text and visual prompts, in customized generation tasks. To explore the core reasons, we decompose SDS into a difference term and a classifier-free guidance term. Our analysis identifies the core issue as arising from the difference term and the random noise addition during the optimization process, both contributing to deviations from the target mode during distillation. To address this, we propose a novel algorithm, Classifier Score Matching (CSM), which removes the difference term in SDS and uses a deterministic noise addition process to reduce noise during optimization, effectively overcoming the low-quality limitations of SDS in our customized generation framework. Based on CSM, we integrate visual prompt information with an attention fusion mechanism and sampling guidance techniques, forming the Visual Prompt CSM (VPCSM) algorithm. Furthermore, we introduce a Semantic-Geometry Calibration (SGC) module to enhance quality through improved textual information integration. We present our approach as TV-3DG, with extensive experiments demonstrating its capability to achieve stable, high-quality, customized 3D generation. Project page: \url{https://yjhboy.github.io/TV-3DG}
format Preprint
id arxiv_https___arxiv_org_abs_2410_21299
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle TV-3DG: Mastering Text-to-3D Customized Generation with Visual Prompt
Yang, Jiahui
Di, Donglin
Ma, Baorui
Yang, Xun
Ma, Yongjia
Sun, Wenzhang
Chen, Wei
Cui, Jianxun
Xue, Zhou
Wang, Meng
Liu, Yebin
Computer Vision and Pattern Recognition
In recent years, advancements in generative models have significantly expanded the capabilities of text-to-3D generation. Many approaches rely on Score Distillation Sampling (SDS) technology. However, SDS struggles to accommodate multi-condition inputs, such as text and visual prompts, in customized generation tasks. To explore the core reasons, we decompose SDS into a difference term and a classifier-free guidance term. Our analysis identifies the core issue as arising from the difference term and the random noise addition during the optimization process, both contributing to deviations from the target mode during distillation. To address this, we propose a novel algorithm, Classifier Score Matching (CSM), which removes the difference term in SDS and uses a deterministic noise addition process to reduce noise during optimization, effectively overcoming the low-quality limitations of SDS in our customized generation framework. Based on CSM, we integrate visual prompt information with an attention fusion mechanism and sampling guidance techniques, forming the Visual Prompt CSM (VPCSM) algorithm. Furthermore, we introduce a Semantic-Geometry Calibration (SGC) module to enhance quality through improved textual information integration. We present our approach as TV-3DG, with extensive experiments demonstrating its capability to achieve stable, high-quality, customized 3D generation. Project page: \url{https://yjhboy.github.io/TV-3DG}
title TV-3DG: Mastering Text-to-3D Customized Generation with Visual Prompt
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2410.21299