ZeroNVS: Zero-Shot 360-Degree View Synthesis from a Single Image
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866913326848540672 |
|---|---|
| author | Sargent, Kyle Li, Zizhang Shah, Tanmay Herrmann, Charles Yu, Hong-Xing Zhang, Yunzhi Chan, Eric Ryan Lagun, Dmitry Fei-Fei, Li Sun, Deqing Wu, Jiajun |
| author_facet | Sargent, Kyle Li, Zizhang Shah, Tanmay Herrmann, Charles Yu, Hong-Xing Zhang, Yunzhi Chan, Eric Ryan Lagun, Dmitry Fei-Fei, Li Sun, Deqing Wu, Jiajun |
| contents | We introduce a 3D-aware diffusion model, ZeroNVS, for single-image novel view synthesis for in-the-wild scenes. While existing methods are designed for single objects with masked backgrounds, we propose new techniques to address challenges introduced by in-the-wild multi-object scenes with complex backgrounds. Specifically, we train a generative prior on a mixture of data sources that capture object-centric, indoor, and outdoor scenes. To address issues from data mixture such as depth-scale ambiguity, we propose a novel camera conditioning parameterization and normalization scheme. Further, we observe that Score Distillation Sampling (SDS) tends to truncate the distribution of complex backgrounds during distillation of 360-degree scenes, and propose "SDS anchoring" to improve the diversity of synthesized novel views. Our model sets a new state-of-the-art result in LPIPS on the DTU dataset in the zero-shot setting, even outperforming methods specifically trained on DTU. We further adapt the challenging Mip-NeRF 360 dataset as a new benchmark for single-image novel view synthesis, and demonstrate strong performance in this setting. Our code and data are at http://kylesargent.github.io/zeronvs/ |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2310_17994 |
| institution | arXiv |
| publishDate | 2023 |
| record_format | arxiv |
| spellingShingle | ZeroNVS: Zero-Shot 360-Degree View Synthesis from a Single Image Sargent, Kyle Li, Zizhang Shah, Tanmay Herrmann, Charles Yu, Hong-Xing Zhang, Yunzhi Chan, Eric Ryan Lagun, Dmitry Fei-Fei, Li Sun, Deqing Wu, Jiajun Computer Vision and Pattern Recognition Graphics We introduce a 3D-aware diffusion model, ZeroNVS, for single-image novel view synthesis for in-the-wild scenes. While existing methods are designed for single objects with masked backgrounds, we propose new techniques to address challenges introduced by in-the-wild multi-object scenes with complex backgrounds. Specifically, we train a generative prior on a mixture of data sources that capture object-centric, indoor, and outdoor scenes. To address issues from data mixture such as depth-scale ambiguity, we propose a novel camera conditioning parameterization and normalization scheme. Further, we observe that Score Distillation Sampling (SDS) tends to truncate the distribution of complex backgrounds during distillation of 360-degree scenes, and propose "SDS anchoring" to improve the diversity of synthesized novel views. Our model sets a new state-of-the-art result in LPIPS on the DTU dataset in the zero-shot setting, even outperforming methods specifically trained on DTU. We further adapt the challenging Mip-NeRF 360 dataset as a new benchmark for single-image novel view synthesis, and demonstrate strong performance in this setting. Our code and data are at http://kylesargent.github.io/zeronvs/ |
| title | ZeroNVS: Zero-Shot 360-Degree View Synthesis from a Single Image |
| topic | Computer Vision and Pattern Recognition Graphics |
| url | https://arxiv.org/abs/2310.17994 |