Sel3DCraft: Interactive Visual Prompts for User-Friendly Text-to-3D Generation
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866912514993815552 |
|---|---|
| author | Xiang, Nan Liang, Tianyi Huang, Haiwen Jiang, Shiqi Huang, Hao Huang, Yifei Chen, Liangyu Wang, Changbo Li, Chenhui |
| author_facet | Xiang, Nan Liang, Tianyi Huang, Haiwen Jiang, Shiqi Huang, Hao Huang, Yifei Chen, Liangyu Wang, Changbo Li, Chenhui |
| contents | Text-to-3D (T23D) generation has transformed digital content creation, yet remains bottlenecked by blind trial-and-error prompting processes that yield unpredictable results. While visual prompt engineering has advanced in text-to-image domains, its application to 3D generation presents unique challenges requiring multi-view consistency evaluation and spatial understanding. We present Sel3DCraft, a visual prompt engineering system for T23D that transforms unstructured exploration into a guided visual process. Our approach introduces three key innovations: a dual-branch structure combining retrieval and generation for diverse candidate exploration; a multi-view hybrid scoring approach that leverages MLLMs with innovative high-level metrics to assess 3D models with human-expert consistency; and a prompt-driven visual analytics suite that enables intuitive defect identification and refinement. Extensive testing and user studies demonstrate that Sel3DCraft surpasses other T23D systems in supporting creativity for designers. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2508_00428 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Sel3DCraft: Interactive Visual Prompts for User-Friendly Text-to-3D Generation Xiang, Nan Liang, Tianyi Huang, Haiwen Jiang, Shiqi Huang, Hao Huang, Yifei Chen, Liangyu Wang, Changbo Li, Chenhui Graphics Human-Computer Interaction Text-to-3D (T23D) generation has transformed digital content creation, yet remains bottlenecked by blind trial-and-error prompting processes that yield unpredictable results. While visual prompt engineering has advanced in text-to-image domains, its application to 3D generation presents unique challenges requiring multi-view consistency evaluation and spatial understanding. We present Sel3DCraft, a visual prompt engineering system for T23D that transforms unstructured exploration into a guided visual process. Our approach introduces three key innovations: a dual-branch structure combining retrieval and generation for diverse candidate exploration; a multi-view hybrid scoring approach that leverages MLLMs with innovative high-level metrics to assess 3D models with human-expert consistency; and a prompt-driven visual analytics suite that enables intuitive defect identification and refinement. Extensive testing and user studies demonstrate that Sel3DCraft surpasses other T23D systems in supporting creativity for designers. |
| title | Sel3DCraft: Interactive Visual Prompts for User-Friendly Text-to-3D Generation |
| topic | Graphics Human-Computer Interaction |
| url | https://arxiv.org/abs/2508.00428 |