CameraCtrl: Enabling Camera Control for Text-to-Video Generation

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: He, Hao, Xu, Yinghao, Guo, Yuwei, Wetzstein, Gordon, Dai, Bo, Li, Hongsheng, Yang, Ceyuan
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909536894320640
author He, Hao
Xu, Yinghao
Guo, Yuwei
Wetzstein, Gordon
Dai, Bo
Li, Hongsheng
Yang, Ceyuan
author_facet He, Hao
Xu, Yinghao
Guo, Yuwei
Wetzstein, Gordon
Dai, Bo
Li, Hongsheng
Yang, Ceyuan
contents Controllability plays a crucial role in video generation, as it allows users to create and edit content more precisely. Existing models, however, lack control of camera pose that serves as a cinematic language to express deeper narrative nuances. To alleviate this issue, we introduce CameraCtrl, enabling accurate camera pose control for video diffusion models. Our approach explores effective camera trajectory parameterization along with a plug-and-play camera pose control module that is trained on top of a video diffusion model, leaving other modules of the base model untouched. Moreover, a comprehensive study on the effect of various training datasets is conducted, suggesting that videos with diverse camera distributions and similar appearance to the base model indeed enhance controllability and generalization. Experimental results demonstrate the effectiveness of CameraCtrl in achieving precise camera control with different video generation models, marking a step forward in the pursuit of dynamic and customized video storytelling from textual and camera pose inputs.
format Preprint
id arxiv_https___arxiv_org_abs_2404_02101
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle CameraCtrl: Enabling Camera Control for Text-to-Video Generation
He, Hao
Xu, Yinghao
Guo, Yuwei
Wetzstein, Gordon
Dai, Bo
Li, Hongsheng
Yang, Ceyuan
Computer Vision and Pattern Recognition
Controllability plays a crucial role in video generation, as it allows users to create and edit content more precisely. Existing models, however, lack control of camera pose that serves as a cinematic language to express deeper narrative nuances. To alleviate this issue, we introduce CameraCtrl, enabling accurate camera pose control for video diffusion models. Our approach explores effective camera trajectory parameterization along with a plug-and-play camera pose control module that is trained on top of a video diffusion model, leaving other modules of the base model untouched. Moreover, a comprehensive study on the effect of various training datasets is conducted, suggesting that videos with diverse camera distributions and similar appearance to the base model indeed enhance controllability and generalization. Experimental results demonstrate the effectiveness of CameraCtrl in achieving precise camera control with different video generation models, marking a step forward in the pursuit of dynamic and customized video storytelling from textual and camera pose inputs.
title CameraCtrl: Enabling Camera Control for Text-to-Video Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2404.02101