Follow-Your-MultiPose: Tuning-Free Multi-Character Text-to-Video Generation via Pose Guidance

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Beiyuan, Ma, Yue, Fu, Chunlei, Song, Xinyang, Sun, Zhenan, Li, Ziqiang
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917879534845952
author Zhang, Beiyuan
Ma, Yue
Fu, Chunlei
Song, Xinyang
Sun, Zhenan
Li, Ziqiang
author_facet Zhang, Beiyuan
Ma, Yue
Fu, Chunlei
Song, Xinyang
Sun, Zhenan
Li, Ziqiang
contents Text-editable and pose-controllable character video generation is a challenging but prevailing topic with practical applications. However, existing approaches mainly focus on single-object video generation with pose guidance, ignoring the realistic situation that multi-character appear concurrently in a scenario. To tackle this, we propose a novel multi-character video generation framework in a tuning-free manner, which is based on the separated text and pose guidance. Specifically, we first extract character masks from the pose sequence to identify the spatial position for each generating character, and then single prompts for each character are obtained with LLMs for precise text guidance. Moreover, the spatial-aligned cross attention and multi-branch control module are proposed to generate fine grained controllable multi-character video. The visualized results of generating video demonstrate the precise controllability of our method for multi-character generation. We also verify the generality of our method by applying it to various personalized T2I models. Moreover, the quantitative results show that our approach achieves superior performance compared with previous works.
format Preprint
id arxiv_https___arxiv_org_abs_2412_16495
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Follow-Your-MultiPose: Tuning-Free Multi-Character Text-to-Video Generation via Pose Guidance
Zhang, Beiyuan
Ma, Yue
Fu, Chunlei
Song, Xinyang
Sun, Zhenan
Li, Ziqiang
Computer Vision and Pattern Recognition
Multimedia
Text-editable and pose-controllable character video generation is a challenging but prevailing topic with practical applications. However, existing approaches mainly focus on single-object video generation with pose guidance, ignoring the realistic situation that multi-character appear concurrently in a scenario. To tackle this, we propose a novel multi-character video generation framework in a tuning-free manner, which is based on the separated text and pose guidance. Specifically, we first extract character masks from the pose sequence to identify the spatial position for each generating character, and then single prompts for each character are obtained with LLMs for precise text guidance. Moreover, the spatial-aligned cross attention and multi-branch control module are proposed to generate fine grained controllable multi-character video. The visualized results of generating video demonstrate the precise controllability of our method for multi-character generation. We also verify the generality of our method by applying it to various personalized T2I models. Moreover, the quantitative results show that our approach achieves superior performance compared with previous works.
title Follow-Your-MultiPose: Tuning-Free Multi-Character Text-to-Video Generation via Pose Guidance
topic Computer Vision and Pattern Recognition
Multimedia
url https://arxiv.org/abs/2412.16495