FloVD: Optical Flow Meets Video Diffusion Model for Enhanced Camera-Controlled Video Synthesis

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jin, Wonjoon, Dai, Qi, Luo, Chong, Baek, Seung-Hwan, Cho, Sunghyun
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917967648784384
author Jin, Wonjoon
Dai, Qi
Luo, Chong
Baek, Seung-Hwan
Cho, Sunghyun
author_facet Jin, Wonjoon
Dai, Qi
Luo, Chong
Baek, Seung-Hwan
Cho, Sunghyun
contents We present FloVD, a novel video diffusion model for camera-controllable video generation. FloVD leverages optical flow to represent the motions of the camera and moving objects. This approach offers two key benefits. Since optical flow can be directly estimated from videos, our approach allows for the use of arbitrary training videos without ground-truth camera parameters. Moreover, as background optical flow encodes 3D correlation across different viewpoints, our method enables detailed camera control by leveraging the background motion. To synthesize natural object motion while supporting detailed camera control, our framework adopts a two-stage video synthesis pipeline consisting of optical flow generation and flow-conditioned video synthesis. Extensive experiments demonstrate the superiority of our method over previous approaches in terms of accurate camera control and natural object motion synthesis.
format Preprint
id arxiv_https___arxiv_org_abs_2502_08244
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle FloVD: Optical Flow Meets Video Diffusion Model for Enhanced Camera-Controlled Video Synthesis
Jin, Wonjoon
Dai, Qi
Luo, Chong
Baek, Seung-Hwan
Cho, Sunghyun
Computer Vision and Pattern Recognition
We present FloVD, a novel video diffusion model for camera-controllable video generation. FloVD leverages optical flow to represent the motions of the camera and moving objects. This approach offers two key benefits. Since optical flow can be directly estimated from videos, our approach allows for the use of arbitrary training videos without ground-truth camera parameters. Moreover, as background optical flow encodes 3D correlation across different viewpoints, our method enables detailed camera control by leveraging the background motion. To synthesize natural object motion while supporting detailed camera control, our framework adopts a two-stage video synthesis pipeline consisting of optical flow generation and flow-conditioned video synthesis. Extensive experiments demonstrate the superiority of our method over previous approaches in terms of accurate camera control and natural object motion synthesis.
title FloVD: Optical Flow Meets Video Diffusion Model for Enhanced Camera-Controlled Video Synthesis
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2502.08244