GS-LRM: Large Reconstruction Model for 3D Gaussian Splatting

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Zhang, Kai, Bi, Sai, Tan, Hao, Xiangli, Yuanbo, Zhao, Nanxuan, Sunkavalli, Kalyan, Xu, Zexiang
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866911860475822080
author Zhang, Kai
Bi, Sai
Tan, Hao
Xiangli, Yuanbo
Zhao, Nanxuan
Sunkavalli, Kalyan
Xu, Zexiang
author_facet Zhang, Kai
Bi, Sai
Tan, Hao
Xiangli, Yuanbo
Zhao, Nanxuan
Sunkavalli, Kalyan
Xu, Zexiang
contents We propose GS-LRM, a scalable large reconstruction model that can predict high-quality 3D Gaussian primitives from 2-4 posed sparse images in 0.23 seconds on single A100 GPU. Our model features a very simple transformer-based architecture; we patchify input posed images, pass the concatenated multi-view image tokens through a sequence of transformer blocks, and decode final per-pixel Gaussian parameters directly from these tokens for differentiable rendering. In contrast to previous LRMs that can only reconstruct objects, by predicting per-pixel Gaussians, GS-LRM naturally handles scenes with large variations in scale and complexity. We show that our model can work on both object and scene captures by training it on Objaverse and RealEstate10K respectively. In both scenarios, the models outperform state-of-the-art baselines by a wide margin. We also demonstrate applications of our model in downstream 3D generation tasks. Our project webpage is available at: https://sai-bi.github.io/project/gs-lrm/ .
format Preprint
id arxiv_https___arxiv_org_abs_2404_19702
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle GS-LRM: Large Reconstruction Model for 3D Gaussian Splatting
Zhang, Kai
Bi, Sai
Tan, Hao
Xiangli, Yuanbo
Zhao, Nanxuan
Sunkavalli, Kalyan
Xu, Zexiang
Computer Vision and Pattern Recognition
We propose GS-LRM, a scalable large reconstruction model that can predict high-quality 3D Gaussian primitives from 2-4 posed sparse images in 0.23 seconds on single A100 GPU. Our model features a very simple transformer-based architecture; we patchify input posed images, pass the concatenated multi-view image tokens through a sequence of transformer blocks, and decode final per-pixel Gaussian parameters directly from these tokens for differentiable rendering. In contrast to previous LRMs that can only reconstruct objects, by predicting per-pixel Gaussians, GS-LRM naturally handles scenes with large variations in scale and complexity. We show that our model can work on both object and scene captures by training it on Objaverse and RealEstate10K respectively. In both scenarios, the models outperform state-of-the-art baselines by a wide margin. We also demonstrate applications of our model in downstream 3D generation tasks. Our project webpage is available at: https://sai-bi.github.io/project/gs-lrm/ .
title GS-LRM: Large Reconstruction Model for 3D Gaussian Splatting
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2404.19702