MotionBooth: Motion-Aware Customized Text-to-Video Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Jianzong, Li, Xiangtai, Zeng, Yanhong, Zhang, Jiangning, Zhou, Qianyu, Li, Yining, Tong, Yunhai, Chen, Kai
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914996610400256
author Wu, Jianzong
Li, Xiangtai
Zeng, Yanhong
Zhang, Jiangning
Zhou, Qianyu
Li, Yining
Tong, Yunhai
Chen, Kai
author_facet Wu, Jianzong
Li, Xiangtai
Zeng, Yanhong
Zhang, Jiangning
Zhou, Qianyu
Li, Yining
Tong, Yunhai
Chen, Kai
contents In this work, we present MotionBooth, an innovative framework designed for animating customized subjects with precise control over both object and camera movements. By leveraging a few images of a specific object, we efficiently fine-tune a text-to-video model to capture the object's shape and attributes accurately. Our approach presents subject region loss and video preservation loss to enhance the subject's learning performance, along with a subject token cross-attention loss to integrate the customized subject with motion control signals. Additionally, we propose training-free techniques for managing subject and camera motions during inference. In particular, we utilize cross-attention map manipulation to govern subject motion and introduce a novel latent shift module for camera movement control as well. MotionBooth excels in preserving the appearance of subjects while simultaneously controlling the motions in generated videos. Extensive quantitative and qualitative evaluations demonstrate the superiority and effectiveness of our method. Our project page is at https://jianzongwu.github.io/projects/motionbooth
format Preprint
id arxiv_https___arxiv_org_abs_2406_17758
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle MotionBooth: Motion-Aware Customized Text-to-Video Generation
Wu, Jianzong
Li, Xiangtai
Zeng, Yanhong
Zhang, Jiangning
Zhou, Qianyu
Li, Yining
Tong, Yunhai
Chen, Kai
Computer Vision and Pattern Recognition
In this work, we present MotionBooth, an innovative framework designed for animating customized subjects with precise control over both object and camera movements. By leveraging a few images of a specific object, we efficiently fine-tune a text-to-video model to capture the object's shape and attributes accurately. Our approach presents subject region loss and video preservation loss to enhance the subject's learning performance, along with a subject token cross-attention loss to integrate the customized subject with motion control signals. Additionally, we propose training-free techniques for managing subject and camera motions during inference. In particular, we utilize cross-attention map manipulation to govern subject motion and introduce a novel latent shift module for camera movement control as well. MotionBooth excels in preserving the appearance of subjects while simultaneously controlling the motions in generated videos. Extensive quantitative and qualitative evaluations demonstrate the superiority and effectiveness of our method. Our project page is at https://jianzongwu.github.io/projects/motionbooth
title MotionBooth: Motion-Aware Customized Text-to-Video Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2406.17758