Dense-Face: Personalized Face Generation Model via Dense Annotation Prediction

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Guo, Xiao, Tran, Manh, Cheng, Jiaxin, Liu, Xiaoming
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917878598467584
author Guo, Xiao
Tran, Manh
Cheng, Jiaxin
Liu, Xiaoming
author_facet Guo, Xiao
Tran, Manh
Cheng, Jiaxin
Liu, Xiaoming
contents The text-to-image (T2I) personalization diffusion model can generate images of the novel concept based on the user input text caption. However, existing T2I personalized methods either require test-time fine-tuning or fail to generate images that align well with the given text caption. In this work, we propose a new T2I personalization diffusion model, Dense-Face, which can generate face images with a consistent identity as the given reference subject and align well with the text caption. Specifically, we introduce a pose-controllable adapter for the high-fidelity image generation while maintaining the text-based editing ability of the pre-trained stable diffusion (SD). Additionally, we use internal features of the SD UNet to predict dense face annotations, enabling the proposed method to gain domain knowledge in face generation. Empirically, our method achieves state-of-the-art or competitive generation performance in image-text alignment, identity preservation, and pose control.
format Preprint
id arxiv_https___arxiv_org_abs_2412_18149
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Dense-Face: Personalized Face Generation Model via Dense Annotation Prediction
Guo, Xiao
Tran, Manh
Cheng, Jiaxin
Liu, Xiaoming
Computer Vision and Pattern Recognition
The text-to-image (T2I) personalization diffusion model can generate images of the novel concept based on the user input text caption. However, existing T2I personalized methods either require test-time fine-tuning or fail to generate images that align well with the given text caption. In this work, we propose a new T2I personalization diffusion model, Dense-Face, which can generate face images with a consistent identity as the given reference subject and align well with the text caption. Specifically, we introduce a pose-controllable adapter for the high-fidelity image generation while maintaining the text-based editing ability of the pre-trained stable diffusion (SD). Additionally, we use internal features of the SD UNet to predict dense face annotations, enabling the proposed method to gain domain knowledge in face generation. Empirically, our method achieves state-of-the-art or competitive generation performance in image-text alignment, identity preservation, and pose control.
title Dense-Face: Personalized Face Generation Model via Dense Annotation Prediction
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.18149