Saved in:
Bibliographic Details
Main Authors: Rauniyar, Aditya, Alama, Omar, Yong, Silong, Sycara, Katia, Scherer, Sebastian
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2501.06431
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915098941980672
author Rauniyar, Aditya
Alama, Omar
Yong, Silong
Sycara, Katia
Scherer, Sebastian
author_facet Rauniyar, Aditya
Alama, Omar
Yong, Silong
Sycara, Katia
Scherer, Sebastian
contents Recent photorealistic Novel View Synthesis (NVS) advances have increasingly gained attention. However, these approaches remain constrained to small indoor scenes. While optimization-based NVS models have attempted to address this, generalizable feed-forward methods, offering significant advantages, remain underexplored. In this work, we train PixelNeRF, a feed-forward NVS model, on the large-scale UrbanScene3D dataset. We propose four training strategies to cluster and train on this dataset, highlighting that performance is hindered by limited view overlap. To address this, we introduce Aug3D, an augmentation technique that leverages reconstructed scenes using traditional Structure-from-Motion (SfM). Aug3D generates well-conditioned novel views through grid and semantic sampling to enhance feed-forward NVS model learning. Our experiments reveal that reducing the number of views per cluster from 20 to 10 improves PSNR by 10%, but the performance remains suboptimal. Aug3D further addresses this by combining the newly generated novel views with the original dataset, demonstrating its effectiveness in improving the model's ability to predict novel views.
format Preprint
id arxiv_https___arxiv_org_abs_2501_06431
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Aug3D: Augmenting large scale outdoor datasets for Generalizable Novel View Synthesis
Rauniyar, Aditya
Alama, Omar
Yong, Silong
Sycara, Katia
Scherer, Sebastian
Computer Vision and Pattern Recognition
Artificial Intelligence
Robotics
Recent photorealistic Novel View Synthesis (NVS) advances have increasingly gained attention. However, these approaches remain constrained to small indoor scenes. While optimization-based NVS models have attempted to address this, generalizable feed-forward methods, offering significant advantages, remain underexplored. In this work, we train PixelNeRF, a feed-forward NVS model, on the large-scale UrbanScene3D dataset. We propose four training strategies to cluster and train on this dataset, highlighting that performance is hindered by limited view overlap. To address this, we introduce Aug3D, an augmentation technique that leverages reconstructed scenes using traditional Structure-from-Motion (SfM). Aug3D generates well-conditioned novel views through grid and semantic sampling to enhance feed-forward NVS model learning. Our experiments reveal that reducing the number of views per cluster from 20 to 10 improves PSNR by 10%, but the performance remains suboptimal. Aug3D further addresses this by combining the newly generated novel views with the original dataset, demonstrating its effectiveness in improving the model's ability to predict novel views.
title Aug3D: Augmenting large scale outdoor datasets for Generalizable Novel View Synthesis
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Robotics
url https://arxiv.org/abs/2501.06431