MegaScenes: Scene-Level View Synthesis at Scale

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tung, Joseph, Chou, Gene, Cai, Ruojin, Yang, Guandao, Zhang, Kai, Wetzstein, Gordon, Hariharan, Bharath, Snavely, Noah
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917755017494528
author Tung, Joseph
Chou, Gene
Cai, Ruojin
Yang, Guandao
Zhang, Kai
Wetzstein, Gordon
Hariharan, Bharath
Snavely, Noah
author_facet Tung, Joseph
Chou, Gene
Cai, Ruojin
Yang, Guandao
Zhang, Kai
Wetzstein, Gordon
Hariharan, Bharath
Snavely, Noah
contents Scene-level novel view synthesis (NVS) is fundamental to many vision and graphics applications. Recently, pose-conditioned diffusion models have led to significant progress by extracting 3D information from 2D foundation models, but these methods are limited by the lack of scene-level training data. Common dataset choices either consist of isolated objects (Objaverse), or of object-centric scenes with limited pose distributions (DTU, CO3D). In this paper, we create a large-scale scene-level dataset from Internet photo collections, called MegaScenes, which contains over 100K structure from motion (SfM) reconstructions from around the world. Internet photos represent a scalable data source but come with challenges such as lighting and transient objects. We address these issues to further create a subset suitable for the task of NVS. Additionally, we analyze failure cases of state-of-the-art NVS methods and significantly improve generation consistency. Through extensive experiments, we validate the effectiveness of both our dataset and method on generating in-the-wild scenes. For details on the dataset and code, see our project page at https://megascenes.github.io.
format Preprint
id arxiv_https___arxiv_org_abs_2406_11819
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle MegaScenes: Scene-Level View Synthesis at Scale
Tung, Joseph
Chou, Gene
Cai, Ruojin
Yang, Guandao
Zhang, Kai
Wetzstein, Gordon
Hariharan, Bharath
Snavely, Noah
Computer Vision and Pattern Recognition
Scene-level novel view synthesis (NVS) is fundamental to many vision and graphics applications. Recently, pose-conditioned diffusion models have led to significant progress by extracting 3D information from 2D foundation models, but these methods are limited by the lack of scene-level training data. Common dataset choices either consist of isolated objects (Objaverse), or of object-centric scenes with limited pose distributions (DTU, CO3D). In this paper, we create a large-scale scene-level dataset from Internet photo collections, called MegaScenes, which contains over 100K structure from motion (SfM) reconstructions from around the world. Internet photos represent a scalable data source but come with challenges such as lighting and transient objects. We address these issues to further create a subset suitable for the task of NVS. Additionally, we analyze failure cases of state-of-the-art NVS methods and significantly improve generation consistency. Through extensive experiments, we validate the effectiveness of both our dataset and method on generating in-the-wild scenes. For details on the dataset and code, see our project page at https://megascenes.github.io.
title MegaScenes: Scene-Level View Synthesis at Scale
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2406.11819