CustomVideo: Customizing Text-to-Video Generation with Multiple Subjects

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Zhao, Li, Aoxue, Zhu, Lingting, Guo, Yong, Dou, Qi, Li, Zhenguo
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908615451869184
author Wang, Zhao
Li, Aoxue
Zhu, Lingting
Guo, Yong
Dou, Qi
Li, Zhenguo
author_facet Wang, Zhao
Li, Aoxue
Zhu, Lingting
Guo, Yong
Dou, Qi
Li, Zhenguo
contents Customized text-to-video generation aims to generate high-quality videos guided by text prompts and subject references. Current approaches for personalizing text-to-video generation suffer from tackling multiple subjects, which is a more challenging and practical scenario. In this work, our aim is to promote multi-subject guided text-to-video customization. We propose CustomVideo, a novel framework that can generate identity-preserving videos with the guidance of multiple subjects. To be specific, firstly, we encourage the co-occurrence of multiple subjects via composing them in a single image. Further, upon a basic text-to-video diffusion model, we design a simple yet effective attention control strategy to disentangle different subjects in the latent space of diffusion model. Moreover, to help the model focus on the specific area of the object, we segment the object from given reference images and provide a corresponding object mask for attention learning. Also, we collect a multi-subject text-to-video generation dataset as a comprehensive benchmark. Extensive qualitative, quantitative, and user study results demonstrate the superiority of our method compared to previous state-of-the-art approaches. The project page is https://kyfafyd.wang/projects/customvideo.
format Preprint
id arxiv_https___arxiv_org_abs_2401_09962
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle CustomVideo: Customizing Text-to-Video Generation with Multiple Subjects
Wang, Zhao
Li, Aoxue
Zhu, Lingting
Guo, Yong
Dou, Qi
Li, Zhenguo
Computer Vision and Pattern Recognition
Customized text-to-video generation aims to generate high-quality videos guided by text prompts and subject references. Current approaches for personalizing text-to-video generation suffer from tackling multiple subjects, which is a more challenging and practical scenario. In this work, our aim is to promote multi-subject guided text-to-video customization. We propose CustomVideo, a novel framework that can generate identity-preserving videos with the guidance of multiple subjects. To be specific, firstly, we encourage the co-occurrence of multiple subjects via composing them in a single image. Further, upon a basic text-to-video diffusion model, we design a simple yet effective attention control strategy to disentangle different subjects in the latent space of diffusion model. Moreover, to help the model focus on the specific area of the object, we segment the object from given reference images and provide a corresponding object mask for attention learning. Also, we collect a multi-subject text-to-video generation dataset as a comprehensive benchmark. Extensive qualitative, quantitative, and user study results demonstrate the superiority of our method compared to previous state-of-the-art approaches. The project page is https://kyfafyd.wang/projects/customvideo.
title CustomVideo: Customizing Text-to-Video Generation with Multiple Subjects
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2401.09962