LRM: Large Reconstruction Model for Single Image to 3D

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Hong, Yicong, Zhang, Kai, Gu, Jiuxiang, Bi, Sai, Zhou, Yang, Liu, Difan, Liu, Feng, Sunkavalli, Kalyan, Bui, Trung, Tan, Hao
Formato: Preprint
Publicado: 2023
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866916154073677824
author Hong, Yicong
Zhang, Kai
Gu, Jiuxiang
Bi, Sai
Zhou, Yang
Liu, Difan
Liu, Feng
Sunkavalli, Kalyan
Bui, Trung
Tan, Hao
author_facet Hong, Yicong
Zhang, Kai
Gu, Jiuxiang
Bi, Sai
Zhou, Yang
Liu, Difan
Liu, Feng
Sunkavalli, Kalyan
Bui, Trung
Tan, Hao
contents We propose the first Large Reconstruction Model (LRM) that predicts the 3D model of an object from a single input image within just 5 seconds. In contrast to many previous methods that are trained on small-scale datasets such as ShapeNet in a category-specific fashion, LRM adopts a highly scalable transformer-based architecture with 500 million learnable parameters to directly predict a neural radiance field (NeRF) from the input image. We train our model in an end-to-end manner on massive multi-view data containing around 1 million objects, including both synthetic renderings from Objaverse and real captures from MVImgNet. This combination of a high-capacity model and large-scale training data empowers our model to be highly generalizable and produce high-quality 3D reconstructions from various testing inputs, including real-world in-the-wild captures and images created by generative models. Video demos and interactable 3D meshes can be found on our LRM project webpage: https://yiconghong.me/LRM.
format Preprint
id arxiv_https___arxiv_org_abs_2311_04400
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle LRM: Large Reconstruction Model for Single Image to 3D
Hong, Yicong
Zhang, Kai
Gu, Jiuxiang
Bi, Sai
Zhou, Yang
Liu, Difan
Liu, Feng
Sunkavalli, Kalyan
Bui, Trung
Tan, Hao
Computer Vision and Pattern Recognition
Artificial Intelligence
Graphics
Machine Learning
We propose the first Large Reconstruction Model (LRM) that predicts the 3D model of an object from a single input image within just 5 seconds. In contrast to many previous methods that are trained on small-scale datasets such as ShapeNet in a category-specific fashion, LRM adopts a highly scalable transformer-based architecture with 500 million learnable parameters to directly predict a neural radiance field (NeRF) from the input image. We train our model in an end-to-end manner on massive multi-view data containing around 1 million objects, including both synthetic renderings from Objaverse and real captures from MVImgNet. This combination of a high-capacity model and large-scale training data empowers our model to be highly generalizable and produce high-quality 3D reconstructions from various testing inputs, including real-world in-the-wild captures and images created by generative models. Video demos and interactable 3D meshes can be found on our LRM project webpage: https://yiconghong.me/LRM.
title LRM: Large Reconstruction Model for Single Image to 3D
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Graphics
Machine Learning
url https://arxiv.org/abs/2311.04400