3DFroMLLM: 3D Prototype Generation only from Pretrained Multimodal LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ahmed, Noor, Braunstein, Cameron, Eger, Steffen, Ilg, Eddy
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915442197528576
author Ahmed, Noor
Braunstein, Cameron
Eger, Steffen
Ilg, Eddy
author_facet Ahmed, Noor
Braunstein, Cameron
Eger, Steffen
Ilg, Eddy
contents Recent Multi-Modal Large Language Models (MLLMs) have demonstrated strong capabilities in learning joint representations from text and images. However, their spatial reasoning remains limited. We introduce 3DFroMLLM, a novel framework that enables the generation of 3D object prototypes directly from MLLMs, including geometry and part labels. Our pipeline is agentic, comprising a designer, coder, and visual inspector operating in a refinement loop. Notably, our approach requires no additional training data or detailed user instructions. Building on prior work in 2D generation, we demonstrate that rendered images produced by our framework can be effectively used for image classification pretraining tasks and outperforms previous methods by 15%. As a compelling real-world use case, we show that the generated prototypes can be leveraged to improve fine-grained vision-language models by using the rendered, part-labeled prototypes to fine-tune CLIP for part segmentation and achieving a 55% accuracy improvement without relying on any additional human-labeled data.
format Preprint
id arxiv_https___arxiv_org_abs_2508_08821
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle 3DFroMLLM: 3D Prototype Generation only from Pretrained Multimodal LLMs
Ahmed, Noor
Braunstein, Cameron
Eger, Steffen
Ilg, Eddy
Computer Vision and Pattern Recognition
Recent Multi-Modal Large Language Models (MLLMs) have demonstrated strong capabilities in learning joint representations from text and images. However, their spatial reasoning remains limited. We introduce 3DFroMLLM, a novel framework that enables the generation of 3D object prototypes directly from MLLMs, including geometry and part labels. Our pipeline is agentic, comprising a designer, coder, and visual inspector operating in a refinement loop. Notably, our approach requires no additional training data or detailed user instructions. Building on prior work in 2D generation, we demonstrate that rendered images produced by our framework can be effectively used for image classification pretraining tasks and outperforms previous methods by 15%. As a compelling real-world use case, we show that the generated prototypes can be leveraged to improve fine-grained vision-language models by using the rendered, part-labeled prototypes to fine-tune CLIP for part segmentation and achieving a 55% accuracy improvement without relying on any additional human-labeled data.
title 3DFroMLLM: 3D Prototype Generation only from Pretrained Multimodal LLMs
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2508.08821