ConsDreamer: Advancing Multi-View Consistency for Zero-Shot Text-to-3D Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhou, Yuan, Jin, Shilong, Hua, Litao, Lv, Wanjun, Duan, Haoran, Han, Jungong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913060348755968
author Zhou, Yuan
Jin, Shilong
Hua, Litao
Lv, Wanjun
Duan, Haoran
Han, Jungong
author_facet Zhou, Yuan
Jin, Shilong
Hua, Litao
Lv, Wanjun
Duan, Haoran
Han, Jungong
contents Recent advances in zero-shot text-to-3D generation have revolutionized 3D content creation by enabling direct synthesis from textual descriptions. While state-of-the-art methods leverage 3D Gaussian Splatting with score distillation to enhance multi-view rendering through pre-trained text-to-image (T2I) models, they suffer from inherent prior view biases in T2I priors. These biases lead to inconsistent 3D generation, particularly manifesting as the multi-face Janus problem, where objects exhibit conflicting features across views. To address this fundamental challenge, we propose ConsDreamer, a novel method that mitigates view bias by refining both the conditional and unconditional terms in the score distillation process: (1) a View Disentanglement Module (VDM) that eliminates viewpoint biases in conditional prompts by decoupling irrelevant view components and injecting precise view control; and (2) a similarity-based partial order loss that enforces geometric consistency in the unconditional term by aligning cosine similarities with azimuth relationships. Extensive experiments demonstrate that ConsDreamer can be seamlessly integrated into various 3D representations and score distillation paradigms, effectively mitigating the multi-face Janus problem.
format Preprint
id arxiv_https___arxiv_org_abs_2504_02316
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ConsDreamer: Advancing Multi-View Consistency for Zero-Shot Text-to-3D Generation
Zhou, Yuan
Jin, Shilong
Hua, Litao
Lv, Wanjun
Duan, Haoran
Han, Jungong
Computer Vision and Pattern Recognition
Artificial Intelligence
Recent advances in zero-shot text-to-3D generation have revolutionized 3D content creation by enabling direct synthesis from textual descriptions. While state-of-the-art methods leverage 3D Gaussian Splatting with score distillation to enhance multi-view rendering through pre-trained text-to-image (T2I) models, they suffer from inherent prior view biases in T2I priors. These biases lead to inconsistent 3D generation, particularly manifesting as the multi-face Janus problem, where objects exhibit conflicting features across views. To address this fundamental challenge, we propose ConsDreamer, a novel method that mitigates view bias by refining both the conditional and unconditional terms in the score distillation process: (1) a View Disentanglement Module (VDM) that eliminates viewpoint biases in conditional prompts by decoupling irrelevant view components and injecting precise view control; and (2) a similarity-based partial order loss that enforces geometric consistency in the unconditional term by aligning cosine similarities with azimuth relationships. Extensive experiments demonstrate that ConsDreamer can be seamlessly integrated into various 3D representations and score distillation paradigms, effectively mitigating the multi-face Janus problem.
title ConsDreamer: Advancing Multi-View Consistency for Zero-Shot Text-to-3D Generation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2504.02316