AssetFormer: Modular 3D Assets Generation with Autoregressive Transformer

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhu, Lingting, Qian, Shengju, Fan, Haidi, Dong, Jiayu, Jin, Zhenchao, Zhou, Siwei, Dong, Gen, Wang, Xin, Yu, Lequan
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911443473924096
author Zhu, Lingting
Qian, Shengju
Fan, Haidi
Dong, Jiayu
Jin, Zhenchao
Zhou, Siwei
Dong, Gen
Wang, Xin
Yu, Lequan
author_facet Zhu, Lingting
Qian, Shengju
Fan, Haidi
Dong, Jiayu
Jin, Zhenchao
Zhou, Siwei
Dong, Gen
Wang, Xin
Yu, Lequan
contents The digital industry demands high-quality, diverse modular 3D assets, especially for user-generated content~(UGC). In this work, we introduce AssetFormer, an autoregressive Transformer-based model designed to generate modular 3D assets from textual descriptions. Our pilot study leverages real-world modular assets collected from online platforms. AssetFormer tackles the challenge of creating assets composed of primitives that adhere to constrained design parameters for various applications. By innovatively adapting module sequencing and decoding techniques inspired by language models, our approach enhances asset generation quality through autoregressive modeling. Initial results indicate the effectiveness of AssetFormer in streamlining asset creation for professional development and UGC scenarios. This work presents a flexible framework extendable to various types of modular 3D assets, contributing to the broader field of 3D content generation. The code is available at https://github.com/Advocate99/AssetFormer.
format Preprint
id arxiv_https___arxiv_org_abs_2602_12100
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle AssetFormer: Modular 3D Assets Generation with Autoregressive Transformer
Zhu, Lingting
Qian, Shengju
Fan, Haidi
Dong, Jiayu
Jin, Zhenchao
Zhou, Siwei
Dong, Gen
Wang, Xin
Yu, Lequan
Computer Vision and Pattern Recognition
The digital industry demands high-quality, diverse modular 3D assets, especially for user-generated content~(UGC). In this work, we introduce AssetFormer, an autoregressive Transformer-based model designed to generate modular 3D assets from textual descriptions. Our pilot study leverages real-world modular assets collected from online platforms. AssetFormer tackles the challenge of creating assets composed of primitives that adhere to constrained design parameters for various applications. By innovatively adapting module sequencing and decoding techniques inspired by language models, our approach enhances asset generation quality through autoregressive modeling. Initial results indicate the effectiveness of AssetFormer in streamlining asset creation for professional development and UGC scenarios. This work presents a flexible framework extendable to various types of modular 3D assets, contributing to the broader field of 3D content generation. The code is available at https://github.com/Advocate99/AssetFormer.
title AssetFormer: Modular 3D Assets Generation with Autoregressive Transformer
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2602.12100