A Generalist FaceX via Learning Unified Facial Representation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Han, Yue, Zhang, Jiangning, Zhu, Junwei, Li, Xiangtai, Ge, Yanhao, Li, Wei, Wang, Chengjie, Liu, Yong, Liu, Xiaoming, Tai, Ying
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913181930094592
author Han, Yue
Zhang, Jiangning
Zhu, Junwei
Li, Xiangtai
Ge, Yanhao
Li, Wei
Wang, Chengjie
Liu, Yong
Liu, Xiaoming
Tai, Ying
author_facet Han, Yue
Zhang, Jiangning
Zhu, Junwei
Li, Xiangtai
Ge, Yanhao
Li, Wei
Wang, Chengjie
Liu, Yong
Liu, Xiaoming
Tai, Ying
contents This work presents FaceX framework, a novel facial generalist model capable of handling diverse facial tasks simultaneously. To achieve this goal, we initially formulate a unified facial representation for a broad spectrum of facial editing tasks, which macroscopically decomposes a face into fundamental identity, intra-personal variation, and environmental factors. Based on this, we introduce Facial Omni-Representation Decomposing (FORD) for seamless manipulation of various facial components, microscopically decomposing the core aspects of most facial editing tasks. Furthermore, by leveraging the prior of a pretrained StableDiffusion (SD) to enhance generation quality and accelerate training, we design Facial Omni-Representation Steering (FORS) to first assemble unified facial representations and then effectively steer the SD-aware generation process by the efficient Facial Representation Controller (FRC). %Without any additional features, Our versatile FaceX achieves competitive performance compared to elaborate task-specific models on popular facial editing tasks. Full codes and models will be available at https://github.com/diffusion-facex/FaceX.
format Preprint
id arxiv_https___arxiv_org_abs_2401_00551
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle A Generalist FaceX via Learning Unified Facial Representation
Han, Yue
Zhang, Jiangning
Zhu, Junwei
Li, Xiangtai
Ge, Yanhao
Li, Wei
Wang, Chengjie
Liu, Yong
Liu, Xiaoming
Tai, Ying
Computer Vision and Pattern Recognition
This work presents FaceX framework, a novel facial generalist model capable of handling diverse facial tasks simultaneously. To achieve this goal, we initially formulate a unified facial representation for a broad spectrum of facial editing tasks, which macroscopically decomposes a face into fundamental identity, intra-personal variation, and environmental factors. Based on this, we introduce Facial Omni-Representation Decomposing (FORD) for seamless manipulation of various facial components, microscopically decomposing the core aspects of most facial editing tasks. Furthermore, by leveraging the prior of a pretrained StableDiffusion (SD) to enhance generation quality and accelerate training, we design Facial Omni-Representation Steering (FORS) to first assemble unified facial representations and then effectively steer the SD-aware generation process by the efficient Facial Representation Controller (FRC). %Without any additional features, Our versatile FaceX achieves competitive performance compared to elaborate task-specific models on popular facial editing tasks. Full codes and models will be available at https://github.com/diffusion-facex/FaceX.
title A Generalist FaceX via Learning Unified Facial Representation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2401.00551