Drive-1-to-3: Enriching Diffusion Priors for Novel View Synthesis of Real Vehicles

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Lin, Chuang, Zhuang, Bingbing, Sun, Shanlin, Jiang, Ziyu, Cai, Jianfei, Chandraker, Manmohan
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915071095996416
author Lin, Chuang
Zhuang, Bingbing
Sun, Shanlin
Jiang, Ziyu
Cai, Jianfei
Chandraker, Manmohan
author_facet Lin, Chuang
Zhuang, Bingbing
Sun, Shanlin
Jiang, Ziyu
Cai, Jianfei
Chandraker, Manmohan
contents The recent advent of large-scale 3D data, e.g. Objaverse, has led to impressive progress in training pose-conditioned diffusion models for novel view synthesis. However, due to the synthetic nature of such 3D data, their performance drops significantly when applied to real-world images. This paper consolidates a set of good practices to finetune large pretrained models for a real-world task -- harvesting vehicle assets for autonomous driving applications. To this end, we delve into the discrepancies between the synthetic data and real driving data, then develop several strategies to account for them properly. Specifically, we start with a virtual camera rotation of real images to ensure geometric alignment with synthetic data and consistency with the pose manifold defined by pretrained models. We also identify important design choices in object-centric data curation to account for varying object distances in real driving scenes -- learn across varying object scales with fixed camera focal length. Further, we perform occlusion-aware training in latent spaces to account for ubiquitous occlusions in real data, and handle large viewpoint changes by leveraging a symmetric prior. Our insights lead to effective finetuning that results in a $68.8\%$ reduction in FID for novel view synthesis over prior arts.
format Preprint
id arxiv_https___arxiv_org_abs_2412_14494
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Drive-1-to-3: Enriching Diffusion Priors for Novel View Synthesis of Real Vehicles
Lin, Chuang
Zhuang, Bingbing
Sun, Shanlin
Jiang, Ziyu
Cai, Jianfei
Chandraker, Manmohan
Computer Vision and Pattern Recognition
The recent advent of large-scale 3D data, e.g. Objaverse, has led to impressive progress in training pose-conditioned diffusion models for novel view synthesis. However, due to the synthetic nature of such 3D data, their performance drops significantly when applied to real-world images. This paper consolidates a set of good practices to finetune large pretrained models for a real-world task -- harvesting vehicle assets for autonomous driving applications. To this end, we delve into the discrepancies between the synthetic data and real driving data, then develop several strategies to account for them properly. Specifically, we start with a virtual camera rotation of real images to ensure geometric alignment with synthetic data and consistency with the pose manifold defined by pretrained models. We also identify important design choices in object-centric data curation to account for varying object distances in real driving scenes -- learn across varying object scales with fixed camera focal length. Further, we perform occlusion-aware training in latent spaces to account for ubiquitous occlusions in real data, and handle large viewpoint changes by leveraging a symmetric prior. Our insights lead to effective finetuning that results in a $68.8\%$ reduction in FID for novel view synthesis over prior arts.
title Drive-1-to-3: Enriching Diffusion Priors for Novel View Synthesis of Real Vehicles
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.14494