MVGamba: Unify 3D Content Generation as State Space Sequence Modeling

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Yi, Xuanyu, Wu, Zike, Shen, Qiuhong, Xu, Qingshan, Zhou, Pan, Lim, Joo-Hwee, Yan, Shuicheng, Wang, Xinchao, Zhang, Hanwang
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910746645889024
author Yi, Xuanyu
Wu, Zike
Shen, Qiuhong
Xu, Qingshan
Zhou, Pan
Lim, Joo-Hwee
Yan, Shuicheng
Wang, Xinchao
Zhang, Hanwang
author_facet Yi, Xuanyu
Wu, Zike
Shen, Qiuhong
Xu, Qingshan
Zhou, Pan
Lim, Joo-Hwee
Yan, Shuicheng
Wang, Xinchao
Zhang, Hanwang
contents Recent 3D large reconstruction models (LRMs) can generate high-quality 3D content in sub-seconds by integrating multi-view diffusion models with scalable multi-view reconstructors. Current works further leverage 3D Gaussian Splatting as 3D representation for improved visual quality and rendering efficiency. However, we observe that existing Gaussian reconstruction models often suffer from multi-view inconsistency and blurred textures. We attribute this to the compromise of multi-view information propagation in favor of adopting powerful yet computationally intensive architectures (e.g., Transformers). To address this issue, we introduce MVGamba, a general and lightweight Gaussian reconstruction model featuring a multi-view Gaussian reconstructor based on the RNN-like State Space Model (SSM). Our Gaussian reconstructor propagates causal context containing multi-view information for cross-view self-refinement while generating a long sequence of Gaussians for fine-detail modeling with linear complexity. With off-the-shelf multi-view diffusion models integrated, MVGamba unifies 3D generation tasks from a single image, sparse images, or text prompts. Extensive experiments demonstrate that MVGamba outperforms state-of-the-art baselines in all 3D content generation scenarios with approximately only $0.1\times$ of the model size.
format Preprint
id arxiv_https___arxiv_org_abs_2406_06367
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle MVGamba: Unify 3D Content Generation as State Space Sequence Modeling
Yi, Xuanyu
Wu, Zike
Shen, Qiuhong
Xu, Qingshan
Zhou, Pan
Lim, Joo-Hwee
Yan, Shuicheng
Wang, Xinchao
Zhang, Hanwang
Computer Vision and Pattern Recognition
Recent 3D large reconstruction models (LRMs) can generate high-quality 3D content in sub-seconds by integrating multi-view diffusion models with scalable multi-view reconstructors. Current works further leverage 3D Gaussian Splatting as 3D representation for improved visual quality and rendering efficiency. However, we observe that existing Gaussian reconstruction models often suffer from multi-view inconsistency and blurred textures. We attribute this to the compromise of multi-view information propagation in favor of adopting powerful yet computationally intensive architectures (e.g., Transformers). To address this issue, we introduce MVGamba, a general and lightweight Gaussian reconstruction model featuring a multi-view Gaussian reconstructor based on the RNN-like State Space Model (SSM). Our Gaussian reconstructor propagates causal context containing multi-view information for cross-view self-refinement while generating a long sequence of Gaussians for fine-detail modeling with linear complexity. With off-the-shelf multi-view diffusion models integrated, MVGamba unifies 3D generation tasks from a single image, sparse images, or text prompts. Extensive experiments demonstrate that MVGamba outperforms state-of-the-art baselines in all 3D content generation scenarios with approximately only $0.1\times$ of the model size.
title MVGamba: Unify 3D Content Generation as State Space Sequence Modeling
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2406.06367