Controlling Space and Time with Diffusion Models

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Watson, Daniel, Saxena, Saurabh, Li, Lala, Tagliasacchi, Andrea, Fleet, David J.
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912335689416704
author Watson, Daniel
Saxena, Saurabh
Li, Lala
Tagliasacchi, Andrea
Fleet, David J.
author_facet Watson, Daniel
Saxena, Saurabh
Li, Lala
Tagliasacchi, Andrea
Fleet, David J.
contents We present 4DiM, a cascaded diffusion model for 4D novel view synthesis (NVS), supporting generation with arbitrary camera trajectories and timestamps, in natural scenes, conditioned on one or more images. With a novel architecture and sampling procedure, we enable training on a mixture of 3D (with camera pose), 4D (pose+time) and video (time but no pose) data, which greatly improves generalization to unseen images and camera pose trajectories over prior works that focus on limited domains (e.g., object centric). 4DiM is the first-ever NVS method with intuitive metric-scale camera pose control enabled by our novel calibration pipeline for structure-from-motion-posed data. Experiments demonstrate that 4DiM outperforms prior 3D NVS models both in terms of image fidelity and pose alignment, while also enabling the generation of scene dynamics. 4DiM provides a general framework for a variety of tasks including single-image-to-3D, two-image-to-video (interpolation and extrapolation), and pose-conditioned video-to-video translation, which we illustrate qualitatively on a variety of scenes. For an overview see https://4d-diffusion.github.io
format Preprint
id arxiv_https___arxiv_org_abs_2407_07860
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Controlling Space and Time with Diffusion Models
Watson, Daniel
Saxena, Saurabh
Li, Lala
Tagliasacchi, Andrea
Fleet, David J.
Computer Vision and Pattern Recognition
We present 4DiM, a cascaded diffusion model for 4D novel view synthesis (NVS), supporting generation with arbitrary camera trajectories and timestamps, in natural scenes, conditioned on one or more images. With a novel architecture and sampling procedure, we enable training on a mixture of 3D (with camera pose), 4D (pose+time) and video (time but no pose) data, which greatly improves generalization to unseen images and camera pose trajectories over prior works that focus on limited domains (e.g., object centric). 4DiM is the first-ever NVS method with intuitive metric-scale camera pose control enabled by our novel calibration pipeline for structure-from-motion-posed data. Experiments demonstrate that 4DiM outperforms prior 3D NVS models both in terms of image fidelity and pose alignment, while also enabling the generation of scene dynamics. 4DiM provides a general framework for a variety of tasks including single-image-to-3D, two-image-to-video (interpolation and extrapolation), and pose-conditioned video-to-video translation, which we illustrate qualitatively on a variety of scenes. For an overview see https://4d-diffusion.github.io
title Controlling Space and Time with Diffusion Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2407.07860