Optical-Flow Guided Prompt Optimization for Coherent Video Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Nam, Hyelin, Kim, Jaemin, Lee, Dohun, Ye, Jong Chul
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913752900698112
author Nam, Hyelin
Kim, Jaemin
Lee, Dohun
Ye, Jong Chul
author_facet Nam, Hyelin
Kim, Jaemin
Lee, Dohun
Ye, Jong Chul
contents While text-to-video diffusion models have made significant strides, many still face challenges in generating videos with temporal consistency. Within diffusion frameworks, guidance techniques have proven effective in enhancing output quality during inference; however, applying these methods to video diffusion models introduces additional complexity of handling computations across entire sequences. To address this, we propose a novel framework called MotionPrompt that guides the video generation process via optical flow. Specifically, we train a discriminator to distinguish optical flow between random pairs of frames from real videos and generated ones. Given that prompts can influence the entire video, we optimize learnable token embeddings during reverse sampling steps by using gradients from a trained discriminator applied to random frame pairs. This approach allows our method to generate visually coherent video sequences that closely reflect natural motion dynamics, without compromising the fidelity of the generated content. We demonstrate the effectiveness of our approach across various models.
format Preprint
id arxiv_https___arxiv_org_abs_2411_15540
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Optical-Flow Guided Prompt Optimization for Coherent Video Generation
Nam, Hyelin
Kim, Jaemin
Lee, Dohun
Ye, Jong Chul
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Image and Video Processing
While text-to-video diffusion models have made significant strides, many still face challenges in generating videos with temporal consistency. Within diffusion frameworks, guidance techniques have proven effective in enhancing output quality during inference; however, applying these methods to video diffusion models introduces additional complexity of handling computations across entire sequences. To address this, we propose a novel framework called MotionPrompt that guides the video generation process via optical flow. Specifically, we train a discriminator to distinguish optical flow between random pairs of frames from real videos and generated ones. Given that prompts can influence the entire video, we optimize learnable token embeddings during reverse sampling steps by using gradients from a trained discriminator applied to random frame pairs. This approach allows our method to generate visually coherent video sequences that closely reflect natural motion dynamics, without compromising the fidelity of the generated content. We demonstrate the effectiveness of our approach across various models.
title Optical-Flow Guided Prompt Optimization for Coherent Video Generation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Image and Video Processing
url https://arxiv.org/abs/2411.15540