CamCtrl3D: Single-Image Scene Exploration with Precise 3D Camera Control

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Popov, Stefan, Raj, Amit, Krainin, Michael, Li, Yuanzhen, Freeman, William T., Rubinstein, Michael
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866910806719856640
author Popov, Stefan
Raj, Amit
Krainin, Michael
Li, Yuanzhen
Freeman, William T.
Rubinstein, Michael
author_facet Popov, Stefan
Raj, Amit
Krainin, Michael
Li, Yuanzhen
Freeman, William T.
Rubinstein, Michael
contents We propose a method for generating fly-through videos of a scene, from a single image and a given camera trajectory. We build upon an image-to-video latent diffusion model. We condition its UNet denoiser on the camera trajectory, using four techniques. (1) We condition the UNet's temporal blocks on raw camera extrinsics, similar to MotionCtrl. (2) We use images containing camera rays and directions, similar to CameraCtrl. (3) We reproject the initial image to subsequent frames and use the resulting video as a condition. (4) We use 2D<=>3D transformers to introduce a global 3D representation, which implicitly conditions on the camera poses. We combine all conditions in a ContolNet-style architecture. We then propose a metric that evaluates overall video quality and the ability to preserve details with view changes, which we use to analyze the trade-offs of individual and combined conditions. Finally, we identify an optimal combination of conditions. We calibrate camera positions in our datasets for scale consistency across scenes, and we train our scene exploration model, CamCtrl3D, demonstrating state-of-theart results.
format Preprint
id arxiv_https___arxiv_org_abs_2501_06006
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CamCtrl3D: Single-Image Scene Exploration with Precise 3D Camera Control
Popov, Stefan
Raj, Amit
Krainin, Michael
Li, Yuanzhen
Freeman, William T.
Rubinstein, Michael
Computer Vision and Pattern Recognition
We propose a method for generating fly-through videos of a scene, from a single image and a given camera trajectory. We build upon an image-to-video latent diffusion model. We condition its UNet denoiser on the camera trajectory, using four techniques. (1) We condition the UNet's temporal blocks on raw camera extrinsics, similar to MotionCtrl. (2) We use images containing camera rays and directions, similar to CameraCtrl. (3) We reproject the initial image to subsequent frames and use the resulting video as a condition. (4) We use 2D<=>3D transformers to introduce a global 3D representation, which implicitly conditions on the camera poses. We combine all conditions in a ContolNet-style architecture. We then propose a metric that evaluates overall video quality and the ability to preserve details with view changes, which we use to analyze the trade-offs of individual and combined conditions. Finally, we identify an optimal combination of conditions. We calibrate camera positions in our datasets for scale consistency across scenes, and we train our scene exploration model, CamCtrl3D, demonstrating state-of-theart results.
title CamCtrl3D: Single-Image Scene Exploration with Precise 3D Camera Control
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2501.06006