StoryMaker: Towards Holistic Consistent Characters in Text-to-image Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhou, Zhengguang, Li, Jing, Li, Huaxia, Chen, Nemo, Tang, Xu
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913508006821888
author Zhou, Zhengguang
Li, Jing
Li, Huaxia
Chen, Nemo
Tang, Xu
author_facet Zhou, Zhengguang
Li, Jing
Li, Huaxia
Chen, Nemo
Tang, Xu
contents Tuning-free personalized image generation methods have achieved significant success in maintaining facial consistency, i.e., identities, even with multiple characters. However, the lack of holistic consistency in scenes with multiple characters hampers these methods' ability to create a cohesive narrative. In this paper, we introduce StoryMaker, a personalization solution that preserves not only facial consistency but also clothing, hairstyles, and body consistency, thus facilitating the creation of a story through a series of images. StoryMaker incorporates conditions based on face identities and cropped character images, which include clothing, hairstyles, and bodies. Specifically, we integrate the facial identity information with the cropped character images using the Positional-aware Perceiver Resampler (PPR) to obtain distinct character features. To prevent intermingling of multiple characters and the background, we separately constrain the cross-attention impact regions of different characters and the background using MSE loss with segmentation masks. Additionally, we train the generation network conditioned on poses to promote decoupling from poses. A LoRA is also employed to enhance fidelity and quality. Experiments underscore the effectiveness of our approach. StoryMaker supports numerous applications and is compatible with other societal plug-ins. Our source codes and model weights are available at https://github.com/RedAIGC/StoryMaker.
format Preprint
id arxiv_https___arxiv_org_abs_2409_12576
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle StoryMaker: Towards Holistic Consistent Characters in Text-to-image Generation
Zhou, Zhengguang
Li, Jing
Li, Huaxia
Chen, Nemo
Tang, Xu
Computer Vision and Pattern Recognition
Tuning-free personalized image generation methods have achieved significant success in maintaining facial consistency, i.e., identities, even with multiple characters. However, the lack of holistic consistency in scenes with multiple characters hampers these methods' ability to create a cohesive narrative. In this paper, we introduce StoryMaker, a personalization solution that preserves not only facial consistency but also clothing, hairstyles, and body consistency, thus facilitating the creation of a story through a series of images. StoryMaker incorporates conditions based on face identities and cropped character images, which include clothing, hairstyles, and bodies. Specifically, we integrate the facial identity information with the cropped character images using the Positional-aware Perceiver Resampler (PPR) to obtain distinct character features. To prevent intermingling of multiple characters and the background, we separately constrain the cross-attention impact regions of different characters and the background using MSE loss with segmentation masks. Additionally, we train the generation network conditioned on poses to promote decoupling from poses. A LoRA is also employed to enhance fidelity and quality. Experiments underscore the effectiveness of our approach. StoryMaker supports numerous applications and is compatible with other societal plug-ins. Our source codes and model weights are available at https://github.com/RedAIGC/StoryMaker.
title StoryMaker: Towards Holistic Consistent Characters in Text-to-image Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2409.12576