NVS-Adapter: Plug-and-Play Novel View Synthesis from a Single Image

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Jeong, Yoonwoo, Lee, Jinwoo, Kim, Chiheon, Cho, Minsu, Lee, Doyup
Format: Preprint
Veröffentlicht: 2023
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909283340255232
author Jeong, Yoonwoo
Lee, Jinwoo
Kim, Chiheon
Cho, Minsu
Lee, Doyup
author_facet Jeong, Yoonwoo
Lee, Jinwoo
Kim, Chiheon
Cho, Minsu
Lee, Doyup
contents Transfer learning of large-scale Text-to-Image (T2I) models has recently shown impressive potential for Novel View Synthesis (NVS) of diverse objects from a single image. While previous methods typically train large models on multi-view datasets for NVS, fine-tuning the whole parameters of T2I models not only demands a high cost but also reduces the generalization capacity of T2I models in generating diverse images in a new domain. In this study, we propose an effective method, dubbed NVS-Adapter, which is a plug-and-play module for a T2I model, to synthesize novel multi-views of visual objects while fully exploiting the generalization capacity of T2I models. NVS-Adapter consists of two main components; view-consistency cross-attention learns the visual correspondences to align the local details of view features, and global semantic conditioning aligns the semantic structure of generated views with the reference view. Experimental results demonstrate that the NVS-Adapter can effectively synthesize geometrically consistent multi-views and also achieve high performance on benchmarks without full fine-tuning of T2I models. The code and data are publicly available in ~\href{https://postech-cvlab.github.io/nvsadapter/}{https://postech-cvlab.github.io/nvsadapter/}.
format Preprint
id arxiv_https___arxiv_org_abs_2312_07315
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle NVS-Adapter: Plug-and-Play Novel View Synthesis from a Single Image
Jeong, Yoonwoo
Lee, Jinwoo
Kim, Chiheon
Cho, Minsu
Lee, Doyup
Computer Vision and Pattern Recognition
Transfer learning of large-scale Text-to-Image (T2I) models has recently shown impressive potential for Novel View Synthesis (NVS) of diverse objects from a single image. While previous methods typically train large models on multi-view datasets for NVS, fine-tuning the whole parameters of T2I models not only demands a high cost but also reduces the generalization capacity of T2I models in generating diverse images in a new domain. In this study, we propose an effective method, dubbed NVS-Adapter, which is a plug-and-play module for a T2I model, to synthesize novel multi-views of visual objects while fully exploiting the generalization capacity of T2I models. NVS-Adapter consists of two main components; view-consistency cross-attention learns the visual correspondences to align the local details of view features, and global semantic conditioning aligns the semantic structure of generated views with the reference view. Experimental results demonstrate that the NVS-Adapter can effectively synthesize geometrically consistent multi-views and also achieve high performance on benchmarks without full fine-tuning of T2I models. The code and data are publicly available in ~\href{https://postech-cvlab.github.io/nvsadapter/}{https://postech-cvlab.github.io/nvsadapter/}.
title NVS-Adapter: Plug-and-Play Novel View Synthesis from a Single Image
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2312.07315