MeshCraft: Exploring Efficient and Controllable Mesh Generation with Flow-based DiTs

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: He, Xianglong, Chen, Junyi, Huang, Di, Liu, Zexiang, Huang, Xiaoshui, Ouyang, Wanli, Yuan, Chun, Li, Yangguang
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910898977767424
author He, Xianglong
Chen, Junyi
Huang, Di
Liu, Zexiang
Huang, Xiaoshui
Ouyang, Wanli
Yuan, Chun
Li, Yangguang
author_facet He, Xianglong
Chen, Junyi
Huang, Di
Liu, Zexiang
Huang, Xiaoshui
Ouyang, Wanli
Yuan, Chun
Li, Yangguang
contents In the domain of 3D content creation, achieving optimal mesh topology through AI models has long been a pursuit for 3D artists. Previous methods, such as MeshGPT, have explored the generation of ready-to-use 3D objects via mesh auto-regressive techniques. While these methods produce visually impressive results, their reliance on token-by-token predictions in the auto-regressive process leads to several significant limitations. These include extremely slow generation speeds and an uncontrollable number of mesh faces. In this paper, we introduce MeshCraft, a novel framework for efficient and controllable mesh generation, which leverages continuous spatial diffusion to generate discrete triangle faces. Specifically, MeshCraft consists of two core components: 1) a transformer-based VAE that encodes raw meshes into continuous face-level tokens and decodes them back to the original meshes, and 2) a flow-based diffusion transformer conditioned on the number of faces, enabling the generation of high-quality 3D meshes with a predefined number of faces. By utilizing the diffusion model for the simultaneous generation of the entire mesh topology, MeshCraft achieves high-fidelity mesh generation at significantly faster speeds compared to auto-regressive methods. Specifically, MeshCraft can generate an 800-face mesh in just 3.2 seconds (35$\times$ faster than existing baselines). Extensive experiments demonstrate that MeshCraft outperforms state-of-the-art techniques in both qualitative and quantitative evaluations on ShapeNet dataset and demonstrates superior performance on Objaverse dataset. Moreover, it integrates seamlessly with existing conditional guidance strategies, showcasing its potential to relieve artists from the time-consuming manual work involved in mesh creation.
format Preprint
id arxiv_https___arxiv_org_abs_2503_23022
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MeshCraft: Exploring Efficient and Controllable Mesh Generation with Flow-based DiTs
He, Xianglong
Chen, Junyi
Huang, Di
Liu, Zexiang
Huang, Xiaoshui
Ouyang, Wanli
Yuan, Chun
Li, Yangguang
Computer Vision and Pattern Recognition
In the domain of 3D content creation, achieving optimal mesh topology through AI models has long been a pursuit for 3D artists. Previous methods, such as MeshGPT, have explored the generation of ready-to-use 3D objects via mesh auto-regressive techniques. While these methods produce visually impressive results, their reliance on token-by-token predictions in the auto-regressive process leads to several significant limitations. These include extremely slow generation speeds and an uncontrollable number of mesh faces. In this paper, we introduce MeshCraft, a novel framework for efficient and controllable mesh generation, which leverages continuous spatial diffusion to generate discrete triangle faces. Specifically, MeshCraft consists of two core components: 1) a transformer-based VAE that encodes raw meshes into continuous face-level tokens and decodes them back to the original meshes, and 2) a flow-based diffusion transformer conditioned on the number of faces, enabling the generation of high-quality 3D meshes with a predefined number of faces. By utilizing the diffusion model for the simultaneous generation of the entire mesh topology, MeshCraft achieves high-fidelity mesh generation at significantly faster speeds compared to auto-regressive methods. Specifically, MeshCraft can generate an 800-face mesh in just 3.2 seconds (35$\times$ faster than existing baselines). Extensive experiments demonstrate that MeshCraft outperforms state-of-the-art techniques in both qualitative and quantitative evaluations on ShapeNet dataset and demonstrates superior performance on Objaverse dataset. Moreover, it integrates seamlessly with existing conditional guidance strategies, showcasing its potential to relieve artists from the time-consuming manual work involved in mesh creation.
title MeshCraft: Exploring Efficient and Controllable Mesh Generation with Flow-based DiTs
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.23022