Saved in:
Bibliographic Details
Main Authors: Bonakdar, Omid, Mozayani, Nasser
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2509.16768
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915504740892672
author Bonakdar, Omid
Mozayani, Nasser
author_facet Bonakdar, Omid
Mozayani, Nasser
contents Generative 3D modeling has advanced rapidly, driven by applications in VR/AR, metaverse, and robotics. However, most methods represent the target object as a closed mesh devoid of any structural information, limiting editing, animation, and semantic understanding. Part-aware 3D generation addresses this problem by decomposing objects into meaningful components, but existing pipelines face challenges: in existing methods, the user has no control over which objects are separated and how model imagine the occluded parts in isolation phase. In this paper, we introduce MMPart, an innovative framework for generating part-aware 3D models from a single image. We first use a VLM to generate a set of prompts based on the input image and user descriptions. In the next step, a generative model generates isolated images of each object based on the initial image and the previous step's prompts as supervisor (which control the pose and guide model how imagine previously occluded areas). Each of those images then enters the multi-view generation stage, where a number of consistent images from different views are generated. Finally, a reconstruction model converts each of these multi-view images into a 3D model.
format Preprint
id arxiv_https___arxiv_org_abs_2509_16768
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MMPart: Harnessing Multi-Modal Large Language Models for Part-Aware 3D Generation
Bonakdar, Omid
Mozayani, Nasser
Computer Vision and Pattern Recognition
Generative 3D modeling has advanced rapidly, driven by applications in VR/AR, metaverse, and robotics. However, most methods represent the target object as a closed mesh devoid of any structural information, limiting editing, animation, and semantic understanding. Part-aware 3D generation addresses this problem by decomposing objects into meaningful components, but existing pipelines face challenges: in existing methods, the user has no control over which objects are separated and how model imagine the occluded parts in isolation phase. In this paper, we introduce MMPart, an innovative framework for generating part-aware 3D models from a single image. We first use a VLM to generate a set of prompts based on the input image and user descriptions. In the next step, a generative model generates isolated images of each object based on the initial image and the previous step's prompts as supervisor (which control the pose and guide model how imagine previously occluded areas). Each of those images then enters the multi-view generation stage, where a number of consistent images from different views are generated. Finally, a reconstruction model converts each of these multi-view images into a 3D model.
title MMPart: Harnessing Multi-Modal Large Language Models for Part-Aware 3D Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2509.16768