MoCA: Mixture-of-Components Attention for Scalable Compositional 3D Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Zhiqi, Li, Wenhuan, Wang, Tengfei, Wang, Zhenwei, Wu, Junta, Wang, Haoyuan, Yang, Yunhan, Huang, Zehuan, Li, Yang, Liu, Peidong, Guo, Chunchao
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917132973899776
author Li, Zhiqi
Li, Wenhuan
Wang, Tengfei
Wang, Zhenwei
Wu, Junta
Wang, Haoyuan
Yang, Yunhan
Huang, Zehuan
Li, Yang
Liu, Peidong
Guo, Chunchao
author_facet Li, Zhiqi
Li, Wenhuan
Wang, Tengfei
Wang, Zhenwei
Wu, Junta
Wang, Haoyuan
Yang, Yunhan
Huang, Zehuan
Li, Yang
Liu, Peidong
Guo, Chunchao
contents Compositionality is critical for 3D object and scene generation, but existing part-aware 3D generation methods suffer from poor scalability due to quadratic global attention costs when increasing the number of components. In this work, we present MoCA, a compositional 3D generative model with two key designs: (1) importance-based component routing that selects top-k relevant components for sparse global attention, and (2) unimportant components compression that preserve contextual priors of unselected components while reducing computational complexity of global attention. With these designs, MoCA enables efficient, fine-grained compositional 3D asset creation with scalable number of components. Extensive experiments show MoCA outperforms baselines on both compositional object and scene generation tasks. Project page: https://lizhiqi49.github.io/MoCA
format Preprint
id arxiv_https___arxiv_org_abs_2512_07628
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MoCA: Mixture-of-Components Attention for Scalable Compositional 3D Generation
Li, Zhiqi
Li, Wenhuan
Wang, Tengfei
Wang, Zhenwei
Wu, Junta
Wang, Haoyuan
Yang, Yunhan
Huang, Zehuan
Li, Yang
Liu, Peidong
Guo, Chunchao
Computer Vision and Pattern Recognition
Compositionality is critical for 3D object and scene generation, but existing part-aware 3D generation methods suffer from poor scalability due to quadratic global attention costs when increasing the number of components. In this work, we present MoCA, a compositional 3D generative model with two key designs: (1) importance-based component routing that selects top-k relevant components for sparse global attention, and (2) unimportant components compression that preserve contextual priors of unselected components while reducing computational complexity of global attention. With these designs, MoCA enables efficient, fine-grained compositional 3D asset creation with scalable number of components. Extensive experiments show MoCA outperforms baselines on both compositional object and scene generation tasks. Project page: https://lizhiqi49.github.io/MoCA
title MoCA: Mixture-of-Components Attention for Scalable Compositional 3D Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.07628