AnimateAnything: Consistent and Controllable Animation for Video Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lei, Guojun, Wang, Chi, Li, Hong, Zhang, Rong, Wang, Yikai, Xu, Weiwei
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912122519158784
author Lei, Guojun
Wang, Chi
Li, Hong
Zhang, Rong
Wang, Yikai
Xu, Weiwei
author_facet Lei, Guojun
Wang, Chi
Li, Hong
Zhang, Rong
Wang, Yikai
Xu, Weiwei
contents We present a unified controllable video generation approach AnimateAnything that facilitates precise and consistent video manipulation across various conditions, including camera trajectories, text prompts, and user motion annotations. Specifically, we carefully design a multi-scale control feature fusion network to construct a common motion representation for different conditions. It explicitly converts all control information into frame-by-frame optical flows. Then we incorporate the optical flows as motion priors to guide final video generation. In addition, to reduce the flickering issues caused by large-scale motion, we propose a frequency-based stabilization module. It can enhance temporal coherence by ensuring the video's frequency domain consistency. Experiments demonstrate that our method outperforms the state-of-the-art approaches. For more details and videos, please refer to the webpage: https://yu-shaonian.github.io/Animate_Anything/.
format Preprint
id arxiv_https___arxiv_org_abs_2411_10836
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle AnimateAnything: Consistent and Controllable Animation for Video Generation
Lei, Guojun
Wang, Chi
Li, Hong
Zhang, Rong
Wang, Yikai
Xu, Weiwei
Computer Vision and Pattern Recognition
We present a unified controllable video generation approach AnimateAnything that facilitates precise and consistent video manipulation across various conditions, including camera trajectories, text prompts, and user motion annotations. Specifically, we carefully design a multi-scale control feature fusion network to construct a common motion representation for different conditions. It explicitly converts all control information into frame-by-frame optical flows. Then we incorporate the optical flows as motion priors to guide final video generation. In addition, to reduce the flickering issues caused by large-scale motion, we propose a frequency-based stabilization module. It can enhance temporal coherence by ensuring the video's frequency domain consistency. Experiments demonstrate that our method outperforms the state-of-the-art approaches. For more details and videos, please refer to the webpage: https://yu-shaonian.github.io/Animate_Anything/.
title AnimateAnything: Consistent and Controllable Animation for Video Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2411.10836