LVT: Large-Scale Scene Reconstruction via Local View Transformers

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Imtiaz, Tooba, Chai, Lucy, Heal, Kathryn, Luo, Xuan, Park, Jungyeon, Dy, Jennifer, Flynn, John
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866916977054842880
author Imtiaz, Tooba
Chai, Lucy
Heal, Kathryn
Luo, Xuan
Park, Jungyeon
Dy, Jennifer
Flynn, John
author_facet Imtiaz, Tooba
Chai, Lucy
Heal, Kathryn
Luo, Xuan
Park, Jungyeon
Dy, Jennifer
Flynn, John
contents Large transformer models are proving to be a powerful tool for 3D vision and novel view synthesis. However, the standard Transformer's well-known quadratic complexity makes it difficult to scale these methods to large scenes. To address this challenge, we propose the Local View Transformer (LVT), a large-scale scene reconstruction and novel view synthesis architecture that circumvents the need for the quadratic attention operation. Motivated by the insight that spatially nearby views provide more useful signal about the local scene composition than distant views, our model processes all information in a local neighborhood around each view. To attend to tokens in nearby views, we leverage a novel positional encoding that conditions on the relative geometric transformation between the query and nearby views. We decode the output of our model into a 3D Gaussian Splat scene representation that includes both color and opacity view-dependence. Taken together, the Local View Transformer enables reconstruction of arbitrarily large, high-resolution scenes in a single forward pass. See our project page for results and interactive demos https://toobaimt.github.io/lvt/.
format Preprint
id arxiv_https___arxiv_org_abs_2509_25001
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LVT: Large-Scale Scene Reconstruction via Local View Transformers
Imtiaz, Tooba
Chai, Lucy
Heal, Kathryn
Luo, Xuan
Park, Jungyeon
Dy, Jennifer
Flynn, John
Computer Vision and Pattern Recognition
Machine Learning
Large transformer models are proving to be a powerful tool for 3D vision and novel view synthesis. However, the standard Transformer's well-known quadratic complexity makes it difficult to scale these methods to large scenes. To address this challenge, we propose the Local View Transformer (LVT), a large-scale scene reconstruction and novel view synthesis architecture that circumvents the need for the quadratic attention operation. Motivated by the insight that spatially nearby views provide more useful signal about the local scene composition than distant views, our model processes all information in a local neighborhood around each view. To attend to tokens in nearby views, we leverage a novel positional encoding that conditions on the relative geometric transformation between the query and nearby views. We decode the output of our model into a 3D Gaussian Splat scene representation that includes both color and opacity view-dependence. Taken together, the Local View Transformer enables reconstruction of arbitrarily large, high-resolution scenes in a single forward pass. See our project page for results and interactive demos https://toobaimt.github.io/lvt/.
title LVT: Large-Scale Scene Reconstruction via Local View Transformers
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2509.25001