CamViG: Camera Aware Image-to-Video Generation with Multimodal Transformers

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Marmon, Andrew, Schindler, Grant, Lezama, José, Kondratyuk, Dan, Seybold, Bryan, Essa, Irfan
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913358518681600
author Marmon, Andrew
Schindler, Grant
Lezama, José
Kondratyuk, Dan
Seybold, Bryan
Essa, Irfan
author_facet Marmon, Andrew
Schindler, Grant
Lezama, José
Kondratyuk, Dan
Seybold, Bryan
Essa, Irfan
contents We extend multimodal transformers to include 3D camera motion as a conditioning signal for the task of video generation. Generative video models are becoming increasingly powerful, thus focusing research efforts on methods of controlling the output of such models. We propose to add virtual 3D camera controls to generative video methods by conditioning generated video on an encoding of three-dimensional camera movement over the course of the generated video. Results demonstrate that we are (1) able to successfully control the camera during video generation, starting from a single frame and a camera signal, and (2) we demonstrate the accuracy of the generated 3D camera paths using traditional computer vision methods.
format Preprint
id arxiv_https___arxiv_org_abs_2405_13195
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle CamViG: Camera Aware Image-to-Video Generation with Multimodal Transformers
Marmon, Andrew
Schindler, Grant
Lezama, José
Kondratyuk, Dan
Seybold, Bryan
Essa, Irfan
Computer Vision and Pattern Recognition
Artificial Intelligence
We extend multimodal transformers to include 3D camera motion as a conditioning signal for the task of video generation. Generative video models are becoming increasingly powerful, thus focusing research efforts on methods of controlling the output of such models. We propose to add virtual 3D camera controls to generative video methods by conditioning generated video on an encoding of three-dimensional camera movement over the course of the generated video. Results demonstrate that we are (1) able to successfully control the camera during video generation, starting from a single frame and a camera signal, and (2) we demonstrate the accuracy of the generated 3D camera paths using traditional computer vision methods.
title CamViG: Camera Aware Image-to-Video Generation with Multimodal Transformers
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2405.13195