OnlyFlow: Optical Flow based Motion Conditioning for Video Diffusion Models
Fuente:
arXiv
Saved in:
| Main Authors: | , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866909876949614592 |
|---|---|
| author | Koroglu, Mathis Caselles-Dupré, Hugo Sanmiguel, Guillaume Jeanneret Cord, Matthieu |
| author_facet | Koroglu, Mathis Caselles-Dupré, Hugo Sanmiguel, Guillaume Jeanneret Cord, Matthieu |
| contents | We consider the problem of text-to-video generation tasks with precise control for various applications such as camera movement control and video-to-video editing. Most methods tacking this problem rely on providing user-defined controls, such as binary masks or camera movement embeddings. In our approach we propose OnlyFlow, an approach leveraging the optical flow firstly extracted from an input video to condition the motion of generated videos. Using a text prompt and an input video, OnlyFlow allows the user to generate videos that respect the motion of the input video as well as the text prompt. This is implemented through an optical flow estimation model applied on the input video, which is then fed to a trainable optical flow encoder. The output feature maps are then injected into the text-to-video backbone model. We perform quantitative, qualitative and user preference studies to show that OnlyFlow positively compares to state-of-the-art methods on a wide range of tasks, even though OnlyFlow was not specifically trained for such tasks. OnlyFlow thus constitutes a versatile, lightweight yet efficient method for controlling motion in text-to-video generation. Models and code will be made available on GitHub and HuggingFace. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2411_10501 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | OnlyFlow: Optical Flow based Motion Conditioning for Video Diffusion Models Koroglu, Mathis Caselles-Dupré, Hugo Sanmiguel, Guillaume Jeanneret Cord, Matthieu Computer Vision and Pattern Recognition Machine Learning I.2.10; I.4.8; I.2.6 We consider the problem of text-to-video generation tasks with precise control for various applications such as camera movement control and video-to-video editing. Most methods tacking this problem rely on providing user-defined controls, such as binary masks or camera movement embeddings. In our approach we propose OnlyFlow, an approach leveraging the optical flow firstly extracted from an input video to condition the motion of generated videos. Using a text prompt and an input video, OnlyFlow allows the user to generate videos that respect the motion of the input video as well as the text prompt. This is implemented through an optical flow estimation model applied on the input video, which is then fed to a trainable optical flow encoder. The output feature maps are then injected into the text-to-video backbone model. We perform quantitative, qualitative and user preference studies to show that OnlyFlow positively compares to state-of-the-art methods on a wide range of tasks, even though OnlyFlow was not specifically trained for such tasks. OnlyFlow thus constitutes a versatile, lightweight yet efficient method for controlling motion in text-to-video generation. Models and code will be made available on GitHub and HuggingFace. |
| title | OnlyFlow: Optical Flow based Motion Conditioning for Video Diffusion Models |
| topic | Computer Vision and Pattern Recognition Machine Learning I.2.10; I.4.8; I.2.6 |
| url | https://arxiv.org/abs/2411.10501 |