StableIdentity: Inserting Anybody into Anywhere at First Sight

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Qinghe, Jia, Xu, Li, Xiaomin, Li, Taiqing, Ma, Liqian, Zhuge, Yunzhi, Lu, Huchuan
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913213385277440
author Wang, Qinghe
Jia, Xu
Li, Xiaomin
Li, Taiqing
Ma, Liqian
Zhuge, Yunzhi
Lu, Huchuan
author_facet Wang, Qinghe
Jia, Xu
Li, Xiaomin
Li, Taiqing
Ma, Liqian
Zhuge, Yunzhi
Lu, Huchuan
contents Recent advances in large pretrained text-to-image models have shown unprecedented capabilities for high-quality human-centric generation, however, customizing face identity is still an intractable problem. Existing methods cannot ensure stable identity preservation and flexible editability, even with several images for each subject during training. In this work, we propose StableIdentity, which allows identity-consistent recontextualization with just one face image. More specifically, we employ a face encoder with an identity prior to encode the input face, and then land the face representation into a space with an editable prior, which is constructed from celeb names. By incorporating identity prior and editability prior, the learned identity can be injected anywhere with various contexts. In addition, we design a masked two-phase diffusion loss to boost the pixel-level perception of the input face and maintain the diversity of generation. Extensive experiments demonstrate our method outperforms previous customization methods. In addition, the learned identity can be flexibly combined with the off-the-shelf modules such as ControlNet. Notably, to the best knowledge, we are the first to directly inject the identity learned from a single image into video/3D generation without finetuning. We believe that the proposed StableIdentity is an important step to unify image, video, and 3D customized generation models.
format Preprint
id arxiv_https___arxiv_org_abs_2401_15975
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle StableIdentity: Inserting Anybody into Anywhere at First Sight
Wang, Qinghe
Jia, Xu
Li, Xiaomin
Li, Taiqing
Ma, Liqian
Zhuge, Yunzhi
Lu, Huchuan
Computer Vision and Pattern Recognition
Recent advances in large pretrained text-to-image models have shown unprecedented capabilities for high-quality human-centric generation, however, customizing face identity is still an intractable problem. Existing methods cannot ensure stable identity preservation and flexible editability, even with several images for each subject during training. In this work, we propose StableIdentity, which allows identity-consistent recontextualization with just one face image. More specifically, we employ a face encoder with an identity prior to encode the input face, and then land the face representation into a space with an editable prior, which is constructed from celeb names. By incorporating identity prior and editability prior, the learned identity can be injected anywhere with various contexts. In addition, we design a masked two-phase diffusion loss to boost the pixel-level perception of the input face and maintain the diversity of generation. Extensive experiments demonstrate our method outperforms previous customization methods. In addition, the learned identity can be flexibly combined with the off-the-shelf modules such as ControlNet. Notably, to the best knowledge, we are the first to directly inject the identity learned from a single image into video/3D generation without finetuning. We believe that the proposed StableIdentity is an important step to unify image, video, and 3D customized generation models.
title StableIdentity: Inserting Anybody into Anywhere at First Sight
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2401.15975