Collaborative Video Diffusion: Consistent Multi-video Generation with Camera Control

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kuang, Zhengfei, Cai, Shengqu, He, Hao, Xu, Yinghao, Li, Hongsheng, Guibas, Leonidas, Wetzstein, Gordon
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929360381935616
author Kuang, Zhengfei
Cai, Shengqu
He, Hao
Xu, Yinghao
Li, Hongsheng
Guibas, Leonidas
Wetzstein, Gordon
author_facet Kuang, Zhengfei
Cai, Shengqu
He, Hao
Xu, Yinghao
Li, Hongsheng
Guibas, Leonidas
Wetzstein, Gordon
contents Research on video generation has recently made tremendous progress, enabling high-quality videos to be generated from text prompts or images. Adding control to the video generation process is an important goal moving forward and recent approaches that condition video generation models on camera trajectories make strides towards it. Yet, it remains challenging to generate a video of the same scene from multiple different camera trajectories. Solutions to this multi-video generation problem could enable large-scale 3D scene generation with editable camera trajectories, among other applications. We introduce collaborative video diffusion (CVD) as an important step towards this vision. The CVD framework includes a novel cross-video synchronization module that promotes consistency between corresponding frames of the same video rendered from different camera poses using an epipolar attention mechanism. Trained on top of a state-of-the-art camera-control module for video generation, CVD generates multiple videos rendered from different camera trajectories with significantly better consistency than baselines, as shown in extensive experiments. Project page: https://collaborativevideodiffusion.github.io/.
format Preprint
id arxiv_https___arxiv_org_abs_2405_17414
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Collaborative Video Diffusion: Consistent Multi-video Generation with Camera Control
Kuang, Zhengfei
Cai, Shengqu
He, Hao
Xu, Yinghao
Li, Hongsheng
Guibas, Leonidas
Wetzstein, Gordon
Computer Vision and Pattern Recognition
Graphics
Research on video generation has recently made tremendous progress, enabling high-quality videos to be generated from text prompts or images. Adding control to the video generation process is an important goal moving forward and recent approaches that condition video generation models on camera trajectories make strides towards it. Yet, it remains challenging to generate a video of the same scene from multiple different camera trajectories. Solutions to this multi-video generation problem could enable large-scale 3D scene generation with editable camera trajectories, among other applications. We introduce collaborative video diffusion (CVD) as an important step towards this vision. The CVD framework includes a novel cross-video synchronization module that promotes consistency between corresponding frames of the same video rendered from different camera poses using an epipolar attention mechanism. Trained on top of a state-of-the-art camera-control module for video generation, CVD generates multiple videos rendered from different camera trajectories with significantly better consistency than baselines, as shown in extensive experiments. Project page: https://collaborativevideodiffusion.github.io/.
title Collaborative Video Diffusion: Consistent Multi-video Generation with Camera Control
topic Computer Vision and Pattern Recognition
Graphics
url https://arxiv.org/abs/2405.17414