Global Latent Neural Rendering

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Tanay, Thomas, Maggioni, Matteo
Format: Preprint
Publié: 2023
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866916151516200960
author Tanay, Thomas
Maggioni, Matteo
author_facet Tanay, Thomas
Maggioni, Matteo
contents A recent trend among generalizable novel view synthesis methods is to learn a rendering operator acting over single camera rays. This approach is promising because it removes the need for explicit volumetric rendering, but it effectively treats target images as collections of independent pixels. Here, we propose to learn a global rendering operator acting over all camera rays jointly. We show that the right representation to enable such rendering is a 5-dimensional plane sweep volume consisting of the projection of the input images on a set of planes facing the target camera. Based on this understanding, we introduce our Convolutional Global Latent Renderer (ConvGLR), an efficient convolutional architecture that performs the rendering operation globally in a low-resolution latent space. Experiments on various datasets under sparse and generalizable setups show that our approach consistently outperforms existing methods by significant margins.
format Preprint
id arxiv_https___arxiv_org_abs_2312_08338
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Global Latent Neural Rendering
Tanay, Thomas
Maggioni, Matteo
Computer Vision and Pattern Recognition
A recent trend among generalizable novel view synthesis methods is to learn a rendering operator acting over single camera rays. This approach is promising because it removes the need for explicit volumetric rendering, but it effectively treats target images as collections of independent pixels. Here, we propose to learn a global rendering operator acting over all camera rays jointly. We show that the right representation to enable such rendering is a 5-dimensional plane sweep volume consisting of the projection of the input images on a set of planes facing the target camera. Based on this understanding, we introduce our Convolutional Global Latent Renderer (ConvGLR), an efficient convolutional architecture that performs the rendering operation globally in a low-resolution latent space. Experiments on various datasets under sparse and generalizable setups show that our approach consistently outperforms existing methods by significant margins.
title Global Latent Neural Rendering
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2312.08338