DreamVideo-2: Zero-Shot Subject-Driven Video Customization with Precise Motion Control

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wei, Yujie, Zhang, Shiwei, Yuan, Hangjie, Wang, Xiang, Qiu, Haonan, Zhao, Rui, Feng, Yutong, Liu, Feng, Huang, Zhizhong, Ye, Jiaxin, Zhang, Yingya, Shan, Hongming
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929548146245632
author Wei, Yujie
Zhang, Shiwei
Yuan, Hangjie
Wang, Xiang
Qiu, Haonan
Zhao, Rui
Feng, Yutong
Liu, Feng
Huang, Zhizhong
Ye, Jiaxin
Zhang, Yingya
Shan, Hongming
author_facet Wei, Yujie
Zhang, Shiwei
Yuan, Hangjie
Wang, Xiang
Qiu, Haonan
Zhao, Rui
Feng, Yutong
Liu, Feng
Huang, Zhizhong
Ye, Jiaxin
Zhang, Yingya
Shan, Hongming
contents Recent advances in customized video generation have enabled users to create videos tailored to both specific subjects and motion trajectories. However, existing methods often require complicated test-time fine-tuning and struggle with balancing subject learning and motion control, limiting their real-world applications. In this paper, we present DreamVideo-2, a zero-shot video customization framework capable of generating videos with a specific subject and motion trajectory, guided by a single image and a bounding box sequence, respectively, and without the need for test-time fine-tuning. Specifically, we introduce reference attention, which leverages the model's inherent capabilities for subject learning, and devise a mask-guided motion module to achieve precise motion control by fully utilizing the robust motion signal of box masks derived from bounding boxes. While these two components achieve their intended functions, we empirically observe that motion control tends to dominate over subject learning. To address this, we propose two key designs: 1) the masked reference attention, which integrates a blended latent mask modeling scheme into reference attention to enhance subject representations at the desired positions, and 2) a reweighted diffusion loss, which differentiates the contributions of regions inside and outside the bounding boxes to ensure a balance between subject and motion control. Extensive experimental results on a newly curated dataset demonstrate that DreamVideo-2 outperforms state-of-the-art methods in both subject customization and motion control. The dataset, code, and models will be made publicly available.
format Preprint
id arxiv_https___arxiv_org_abs_2410_13830
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle DreamVideo-2: Zero-Shot Subject-Driven Video Customization with Precise Motion Control
Wei, Yujie
Zhang, Shiwei
Yuan, Hangjie
Wang, Xiang
Qiu, Haonan
Zhao, Rui
Feng, Yutong
Liu, Feng
Huang, Zhizhong
Ye, Jiaxin
Zhang, Yingya
Shan, Hongming
Computer Vision and Pattern Recognition
Recent advances in customized video generation have enabled users to create videos tailored to both specific subjects and motion trajectories. However, existing methods often require complicated test-time fine-tuning and struggle with balancing subject learning and motion control, limiting their real-world applications. In this paper, we present DreamVideo-2, a zero-shot video customization framework capable of generating videos with a specific subject and motion trajectory, guided by a single image and a bounding box sequence, respectively, and without the need for test-time fine-tuning. Specifically, we introduce reference attention, which leverages the model's inherent capabilities for subject learning, and devise a mask-guided motion module to achieve precise motion control by fully utilizing the robust motion signal of box masks derived from bounding boxes. While these two components achieve their intended functions, we empirically observe that motion control tends to dominate over subject learning. To address this, we propose two key designs: 1) the masked reference attention, which integrates a blended latent mask modeling scheme into reference attention to enhance subject representations at the desired positions, and 2) a reweighted diffusion loss, which differentiates the contributions of regions inside and outside the bounding boxes to ensure a balance between subject and motion control. Extensive experimental results on a newly curated dataset demonstrate that DreamVideo-2 outperforms state-of-the-art methods in both subject customization and motion control. The dataset, code, and models will be made publicly available.
title DreamVideo-2: Zero-Shot Subject-Driven Video Customization with Precise Motion Control
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2410.13830