DisenStudio: Customized Multi-subject Text-to-Video Generation with Disentangled Spatial Control

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Chen, Hong, Wang, Xin, Zhang, Yipeng, Zhou, Yuwei, Zhang, Zeyang, Tang, Siao, Zhu, Wenwu
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866914804001669120
author Chen, Hong
Wang, Xin
Zhang, Yipeng
Zhou, Yuwei
Zhang, Zeyang
Tang, Siao
Zhu, Wenwu
author_facet Chen, Hong
Wang, Xin
Zhang, Yipeng
Zhou, Yuwei
Zhang, Zeyang
Tang, Siao
Zhu, Wenwu
contents Generating customized content in videos has received increasing attention recently. However, existing works primarily focus on customized text-to-video generation for single subject, suffering from subject-missing and attribute-binding problems when the video is expected to contain multiple subjects. Furthermore, existing models struggle to assign the desired actions to the corresponding subjects (action-binding problem), failing to achieve satisfactory multi-subject generation performance. To tackle the problems, in this paper, we propose DisenStudio, a novel framework that can generate text-guided videos for customized multiple subjects, given few images for each subject. Specifically, DisenStudio enhances a pretrained diffusion-based text-to-video model with our proposed spatial-disentangled cross-attention mechanism to associate each subject with the desired action. Then the model is customized for the multiple subjects with the proposed motion-preserved disentangled finetuning, which involves three tuning strategies: multi-subject co-occurrence tuning, masked single-subject tuning, and multi-subject motion-preserved tuning. The first two strategies guarantee the subject occurrence and preserve their visual attributes, and the third strategy helps the model maintain the temporal motion-generation ability when finetuning on static images. We conduct extensive experiments to demonstrate our proposed DisenStudio significantly outperforms existing methods in various metrics. Additionally, we show that DisenStudio can be used as a powerful tool for various controllable generation applications.
format Preprint
id arxiv_https___arxiv_org_abs_2405_12796
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle DisenStudio: Customized Multi-subject Text-to-Video Generation with Disentangled Spatial Control
Chen, Hong
Wang, Xin
Zhang, Yipeng
Zhou, Yuwei
Zhang, Zeyang
Tang, Siao
Zhu, Wenwu
Computer Vision and Pattern Recognition
Generating customized content in videos has received increasing attention recently. However, existing works primarily focus on customized text-to-video generation for single subject, suffering from subject-missing and attribute-binding problems when the video is expected to contain multiple subjects. Furthermore, existing models struggle to assign the desired actions to the corresponding subjects (action-binding problem), failing to achieve satisfactory multi-subject generation performance. To tackle the problems, in this paper, we propose DisenStudio, a novel framework that can generate text-guided videos for customized multiple subjects, given few images for each subject. Specifically, DisenStudio enhances a pretrained diffusion-based text-to-video model with our proposed spatial-disentangled cross-attention mechanism to associate each subject with the desired action. Then the model is customized for the multiple subjects with the proposed motion-preserved disentangled finetuning, which involves three tuning strategies: multi-subject co-occurrence tuning, masked single-subject tuning, and multi-subject motion-preserved tuning. The first two strategies guarantee the subject occurrence and preserve their visual attributes, and the third strategy helps the model maintain the temporal motion-generation ability when finetuning on static images. We conduct extensive experiments to demonstrate our proposed DisenStudio significantly outperforms existing methods in various metrics. Additionally, we show that DisenStudio can be used as a powerful tool for various controllable generation applications.
title DisenStudio: Customized Multi-subject Text-to-Video Generation with Disentangled Spatial Control
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2405.12796