Saved in:
Bibliographic Details
Main Authors: Zhang, Qing, Tong, Jinguang, Zhang, Jing, Hong, Jie, Li, Xuesong
Format: Preprint
Published: 2026
Subjects:
Online Access:https://arxiv.org/abs/2605.19876
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911698500190208
author Zhang, Qing
Tong, Jinguang
Zhang, Jing
Hong, Jie
Li, Xuesong
author_facet Zhang, Qing
Tong, Jinguang
Zhang, Jing
Hong, Jie
Li, Xuesong
contents Text-to-3D generation based on diffusion models often suffers from the Janus problem, leading to inconsistent geometry across viewpoints. This work identifies viewpoint bias in 2D diffusion priors as the main cause and proposes Structural Energy-Guided Sampling (SEGS), a training-free and plug-and-play framework to improve multi-view consistency. SEGS constructs a structural energy in the PCA subspace of U-Net features and injects its gradient into the denoising process. It can be easily integrated into SDS/VSD pipelines without retraining. Experiments show that SEGS reduces the Janus Rate by about 10% on average and improves View-CS scores across multiple baselines, including DreamFusion, Magic3D, and LucidDreamer. This method effectively alleviates viewpoint artifacts while preserving appearance fidelity, providing a flexible solution for high-quality text-to-3D content generation.
format Preprint
id arxiv_https___arxiv_org_abs_2605_19876
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Structural Energy Guidance for View-Consistent Text-to-3D Generation
Zhang, Qing
Tong, Jinguang
Zhang, Jing
Hong, Jie
Li, Xuesong
Computer Vision and Pattern Recognition
Text-to-3D generation based on diffusion models often suffers from the Janus problem, leading to inconsistent geometry across viewpoints. This work identifies viewpoint bias in 2D diffusion priors as the main cause and proposes Structural Energy-Guided Sampling (SEGS), a training-free and plug-and-play framework to improve multi-view consistency. SEGS constructs a structural energy in the PCA subspace of U-Net features and injects its gradient into the denoising process. It can be easily integrated into SDS/VSD pipelines without retraining. Experiments show that SEGS reduces the Janus Rate by about 10% on average and improves View-CS scores across multiple baselines, including DreamFusion, Magic3D, and LucidDreamer. This method effectively alleviates viewpoint artifacts while preserving appearance fidelity, providing a flexible solution for high-quality text-to-3D content generation.
title Structural Energy Guidance for View-Consistent Text-to-3D Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.19876