Sel3DCraft: Interactive Visual Prompts for User-Friendly Text-to-3D Generation

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Xiang, Nan, Liang, Tianyi, Huang, Haiwen, Jiang, Shiqi, Huang, Hao, Huang, Yifei, Chen, Liangyu, Wang, Changbo, Li, Chenhui
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866912514993815552
author Xiang, Nan
Liang, Tianyi
Huang, Haiwen
Jiang, Shiqi
Huang, Hao
Huang, Yifei
Chen, Liangyu
Wang, Changbo
Li, Chenhui
author_facet Xiang, Nan
Liang, Tianyi
Huang, Haiwen
Jiang, Shiqi
Huang, Hao
Huang, Yifei
Chen, Liangyu
Wang, Changbo
Li, Chenhui
contents Text-to-3D (T23D) generation has transformed digital content creation, yet remains bottlenecked by blind trial-and-error prompting processes that yield unpredictable results. While visual prompt engineering has advanced in text-to-image domains, its application to 3D generation presents unique challenges requiring multi-view consistency evaluation and spatial understanding. We present Sel3DCraft, a visual prompt engineering system for T23D that transforms unstructured exploration into a guided visual process. Our approach introduces three key innovations: a dual-branch structure combining retrieval and generation for diverse candidate exploration; a multi-view hybrid scoring approach that leverages MLLMs with innovative high-level metrics to assess 3D models with human-expert consistency; and a prompt-driven visual analytics suite that enables intuitive defect identification and refinement. Extensive testing and user studies demonstrate that Sel3DCraft surpasses other T23D systems in supporting creativity for designers.
format Preprint
id arxiv_https___arxiv_org_abs_2508_00428
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Sel3DCraft: Interactive Visual Prompts for User-Friendly Text-to-3D Generation
Xiang, Nan
Liang, Tianyi
Huang, Haiwen
Jiang, Shiqi
Huang, Hao
Huang, Yifei
Chen, Liangyu
Wang, Changbo
Li, Chenhui
Graphics
Human-Computer Interaction
Text-to-3D (T23D) generation has transformed digital content creation, yet remains bottlenecked by blind trial-and-error prompting processes that yield unpredictable results. While visual prompt engineering has advanced in text-to-image domains, its application to 3D generation presents unique challenges requiring multi-view consistency evaluation and spatial understanding. We present Sel3DCraft, a visual prompt engineering system for T23D that transforms unstructured exploration into a guided visual process. Our approach introduces three key innovations: a dual-branch structure combining retrieval and generation for diverse candidate exploration; a multi-view hybrid scoring approach that leverages MLLMs with innovative high-level metrics to assess 3D models with human-expert consistency; and a prompt-driven visual analytics suite that enables intuitive defect identification and refinement. Extensive testing and user studies demonstrate that Sel3DCraft surpasses other T23D systems in supporting creativity for designers.
title Sel3DCraft: Interactive Visual Prompts for User-Friendly Text-to-3D Generation
topic Graphics
Human-Computer Interaction
url https://arxiv.org/abs/2508.00428