latentSplat: Autoencoding Variational Gaussians for Fast Generalizable 3D Reconstruction
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866910546312298496 |
|---|---|
| author | Wewer, Christopher Raj, Kevin Ilg, Eddy Schiele, Bernt Lenssen, Jan Eric |
| author_facet | Wewer, Christopher Raj, Kevin Ilg, Eddy Schiele, Bernt Lenssen, Jan Eric |
| contents | We present latentSplat, a method to predict semantic Gaussians in a 3D latent space that can be splatted and decoded by a light-weight generative 2D architecture. Existing methods for generalizable 3D reconstruction either do not scale to large scenes and resolutions, or are limited to interpolation of close input views. latentSplat combines the strengths of regression-based and generative approaches while being trained purely on readily available real video data. The core of our method are variational 3D Gaussians, a representation that efficiently encodes varying uncertainty within a latent space consisting of 3D feature Gaussians. From these Gaussians, specific instances can be sampled and rendered via efficient splatting and a fast, generative decoder. We show that latentSplat outperforms previous works in reconstruction quality and generalization, while being fast and scalable to high-resolution data. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2403_16292 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | latentSplat: Autoencoding Variational Gaussians for Fast Generalizable 3D Reconstruction Wewer, Christopher Raj, Kevin Ilg, Eddy Schiele, Bernt Lenssen, Jan Eric Computer Vision and Pattern Recognition We present latentSplat, a method to predict semantic Gaussians in a 3D latent space that can be splatted and decoded by a light-weight generative 2D architecture. Existing methods for generalizable 3D reconstruction either do not scale to large scenes and resolutions, or are limited to interpolation of close input views. latentSplat combines the strengths of regression-based and generative approaches while being trained purely on readily available real video data. The core of our method are variational 3D Gaussians, a representation that efficiently encodes varying uncertainty within a latent space consisting of 3D feature Gaussians. From these Gaussians, specific instances can be sampled and rendered via efficient splatting and a fast, generative decoder. We show that latentSplat outperforms previous works in reconstruction quality and generalization, while being fast and scalable to high-resolution data. |
| title | latentSplat: Autoencoding Variational Gaussians for Fast Generalizable 3D Reconstruction |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2403.16292 |