OnlyFlow: Optical Flow based Motion Conditioning for Video Diffusion Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Koroglu, Mathis, Caselles-Dupré, Hugo, Sanmiguel, Guillaume Jeanneret, Cord, Matthieu
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909876949614592
author Koroglu, Mathis
Caselles-Dupré, Hugo
Sanmiguel, Guillaume Jeanneret
Cord, Matthieu
author_facet Koroglu, Mathis
Caselles-Dupré, Hugo
Sanmiguel, Guillaume Jeanneret
Cord, Matthieu
contents We consider the problem of text-to-video generation tasks with precise control for various applications such as camera movement control and video-to-video editing. Most methods tacking this problem rely on providing user-defined controls, such as binary masks or camera movement embeddings. In our approach we propose OnlyFlow, an approach leveraging the optical flow firstly extracted from an input video to condition the motion of generated videos. Using a text prompt and an input video, OnlyFlow allows the user to generate videos that respect the motion of the input video as well as the text prompt. This is implemented through an optical flow estimation model applied on the input video, which is then fed to a trainable optical flow encoder. The output feature maps are then injected into the text-to-video backbone model. We perform quantitative, qualitative and user preference studies to show that OnlyFlow positively compares to state-of-the-art methods on a wide range of tasks, even though OnlyFlow was not specifically trained for such tasks. OnlyFlow thus constitutes a versatile, lightweight yet efficient method for controlling motion in text-to-video generation. Models and code will be made available on GitHub and HuggingFace.
format Preprint
id arxiv_https___arxiv_org_abs_2411_10501
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle OnlyFlow: Optical Flow based Motion Conditioning for Video Diffusion Models
Koroglu, Mathis
Caselles-Dupré, Hugo
Sanmiguel, Guillaume Jeanneret
Cord, Matthieu
Computer Vision and Pattern Recognition
Machine Learning
I.2.10; I.4.8; I.2.6
We consider the problem of text-to-video generation tasks with precise control for various applications such as camera movement control and video-to-video editing. Most methods tacking this problem rely on providing user-defined controls, such as binary masks or camera movement embeddings. In our approach we propose OnlyFlow, an approach leveraging the optical flow firstly extracted from an input video to condition the motion of generated videos. Using a text prompt and an input video, OnlyFlow allows the user to generate videos that respect the motion of the input video as well as the text prompt. This is implemented through an optical flow estimation model applied on the input video, which is then fed to a trainable optical flow encoder. The output feature maps are then injected into the text-to-video backbone model. We perform quantitative, qualitative and user preference studies to show that OnlyFlow positively compares to state-of-the-art methods on a wide range of tasks, even though OnlyFlow was not specifically trained for such tasks. OnlyFlow thus constitutes a versatile, lightweight yet efficient method for controlling motion in text-to-video generation. Models and code will be made available on GitHub and HuggingFace.
title OnlyFlow: Optical Flow based Motion Conditioning for Video Diffusion Models
topic Computer Vision and Pattern Recognition
Machine Learning
I.2.10; I.4.8; I.2.6
url https://arxiv.org/abs/2411.10501