Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Zhang, Qing, Tong, Jinguang, Zhang, Jing, Hong, Jie, Li, Xuesong
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:https://arxiv.org/abs/2412.02287
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912536587141120
author Zhang, Qing
Tong, Jinguang
Zhang, Jing
Hong, Jie
Li, Xuesong
author_facet Zhang, Qing
Tong, Jinguang
Zhang, Jing
Hong, Jie
Li, Xuesong
contents Despite recent advances in text-to-3D generation techniques, current methods often suffer from geometric inconsistencies, commonly referred to as the Janus Problem. This paper identifies the root cause of the Janus Problem: viewpoint generation bias in diffusion models, which creates a significant gap between the actual generated viewpoint and the expected one required for optimizing the 3D model. To address this issue, we propose a tuning-free approach called the Attention and CLIP Guidance (ACG) mechanism. ACG enhances desired viewpoints by adaptively controlling cross-attention maps, employs CLIP-based view-text similarities to filter out erroneous viewpoints, and uses a coarse-to-fine optimization strategy with staged prompts to progressively refine 3D generation. Extensive experiments demonstrate that our method significantly reduces the Janus Problem without compromising generation speed, establishing ACG as an efficient, plug-and-play component for existing text-to-3D frameworks.
format Preprint
id arxiv_https___arxiv_org_abs_2412_02287
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Improving Viewpoint Consistency in 3D Generation via Structure Feature and CLIP Guidance
Zhang, Qing
Tong, Jinguang
Zhang, Jing
Hong, Jie
Li, Xuesong
Computer Vision and Pattern Recognition
Despite recent advances in text-to-3D generation techniques, current methods often suffer from geometric inconsistencies, commonly referred to as the Janus Problem. This paper identifies the root cause of the Janus Problem: viewpoint generation bias in diffusion models, which creates a significant gap between the actual generated viewpoint and the expected one required for optimizing the 3D model. To address this issue, we propose a tuning-free approach called the Attention and CLIP Guidance (ACG) mechanism. ACG enhances desired viewpoints by adaptively controlling cross-attention maps, employs CLIP-based view-text similarities to filter out erroneous viewpoints, and uses a coarse-to-fine optimization strategy with staged prompts to progressively refine 3D generation. Extensive experiments demonstrate that our method significantly reduces the Janus Problem without compromising generation speed, establishing ACG as an efficient, plug-and-play component for existing text-to-3D frameworks.
title Improving Viewpoint Consistency in 3D Generation via Structure Feature and CLIP Guidance
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.02287