MeshLLM: Empowering Large Language Models to Progressively Understand and Generate 3D Mesh

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Fang, Shuangkang, Shen, I-Chao, Wang, Yufeng, Tsai, Yi-Hsuan, Yang, Yi, Zhou, Shuchang, Ding, Wenrui, Igarashi, Takeo, Yang, Ming-Hsuan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908479154814976
author Fang, Shuangkang
Shen, I-Chao
Wang, Yufeng
Tsai, Yi-Hsuan
Yang, Yi
Zhou, Shuchang
Ding, Wenrui
Igarashi, Takeo
Yang, Ming-Hsuan
author_facet Fang, Shuangkang
Shen, I-Chao
Wang, Yufeng
Tsai, Yi-Hsuan
Yang, Yi
Zhou, Shuchang
Ding, Wenrui
Igarashi, Takeo
Yang, Ming-Hsuan
contents We present MeshLLM, a novel framework that leverages large language models (LLMs) to understand and generate text-serialized 3D meshes. Our approach addresses key limitations in existing methods, including the limited dataset scale when catering to LLMs' token length and the loss of 3D structural information during mesh serialization. We introduce a Primitive-Mesh decomposition strategy, which divides 3D meshes into structurally meaningful subunits. This enables the creation of a large-scale dataset with 1500k+ samples, almost 50 times larger than previous methods, which aligns better with the LLM scaling law principles. Furthermore, we propose inferring face connectivity from vertices and local mesh assembly training strategies, significantly enhancing the LLMs' ability to capture mesh topology and spatial structures. Experiments show that MeshLLM outperforms the state-of-the-art LLaMA-Mesh in both mesh generation quality and shape understanding, highlighting its great potential in processing text-serialized 3D meshes.
format Preprint
id arxiv_https___arxiv_org_abs_2508_01242
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MeshLLM: Empowering Large Language Models to Progressively Understand and Generate 3D Mesh
Fang, Shuangkang
Shen, I-Chao
Wang, Yufeng
Tsai, Yi-Hsuan
Yang, Yi
Zhou, Shuchang
Ding, Wenrui
Igarashi, Takeo
Yang, Ming-Hsuan
Graphics
Computer Vision and Pattern Recognition
We present MeshLLM, a novel framework that leverages large language models (LLMs) to understand and generate text-serialized 3D meshes. Our approach addresses key limitations in existing methods, including the limited dataset scale when catering to LLMs' token length and the loss of 3D structural information during mesh serialization. We introduce a Primitive-Mesh decomposition strategy, which divides 3D meshes into structurally meaningful subunits. This enables the creation of a large-scale dataset with 1500k+ samples, almost 50 times larger than previous methods, which aligns better with the LLM scaling law principles. Furthermore, we propose inferring face connectivity from vertices and local mesh assembly training strategies, significantly enhancing the LLMs' ability to capture mesh topology and spatial structures. Experiments show that MeshLLM outperforms the state-of-the-art LLaMA-Mesh in both mesh generation quality and shape understanding, highlighting its great potential in processing text-serialized 3D meshes.
title MeshLLM: Empowering Large Language Models to Progressively Understand and Generate 3D Mesh
topic Graphics
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2508.01242