Towards Consistent and Controllable Image Synthesis for Face Editing

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wei, Mengting, Varanka, Tuomas, Li, Yante, Jiang, Xingxun, Khor, Huai-Qian, Zhao, Guoying
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914171585560576
author Wei, Mengting
Varanka, Tuomas
Li, Yante
Jiang, Xingxun
Khor, Huai-Qian
Zhao, Guoying
author_facet Wei, Mengting
Varanka, Tuomas
Li, Yante
Jiang, Xingxun
Khor, Huai-Qian
Zhao, Guoying
contents Face editing methods, essential for tasks like virtual avatars, digital human synthesis and identity preservation, have traditionally been built upon GAN-based techniques, while recent focus has shifted to diffusion-based models due to their success in image reconstruction. However, diffusion models still face challenges in controlling specific attributes and preserving the consistency of other unchanged attributes especially the identity characteristics. To address these issues and facilitate more convenient editing of face images, we propose a novel approach that leverages the power of Stable-Diffusion (SD) models and crude 3D face models to control the lighting, facial expression and head pose of a portrait photo. We observe that this task essentially involves the combinations of target background, identity and face attributes aimed to edit. We strive to sufficiently disentangle the control of these factors to enable consistency of face editing. Specifically, our method, coined as RigFace, contains: 1) A Spatial Attribute Encoder that provides presise and decoupled conditions of background, pose, expression and lighting; 2) A high-consistency FaceFusion method that transfers identity features from the Identity Encoder to the denoising UNet of a pre-trained SD model; 3) An Attribute Rigger that injects those conditions into the denoising UNet. Our model achieves comparable or even superior performance in both identity preservation and photorealism compared to existing face editing models.
format Preprint
id arxiv_https___arxiv_org_abs_2502_02465
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Towards Consistent and Controllable Image Synthesis for Face Editing
Wei, Mengting
Varanka, Tuomas
Li, Yante
Jiang, Xingxun
Khor, Huai-Qian
Zhao, Guoying
Computer Vision and Pattern Recognition
Face editing methods, essential for tasks like virtual avatars, digital human synthesis and identity preservation, have traditionally been built upon GAN-based techniques, while recent focus has shifted to diffusion-based models due to their success in image reconstruction. However, diffusion models still face challenges in controlling specific attributes and preserving the consistency of other unchanged attributes especially the identity characteristics. To address these issues and facilitate more convenient editing of face images, we propose a novel approach that leverages the power of Stable-Diffusion (SD) models and crude 3D face models to control the lighting, facial expression and head pose of a portrait photo. We observe that this task essentially involves the combinations of target background, identity and face attributes aimed to edit. We strive to sufficiently disentangle the control of these factors to enable consistency of face editing. Specifically, our method, coined as RigFace, contains: 1) A Spatial Attribute Encoder that provides presise and decoupled conditions of background, pose, expression and lighting; 2) A high-consistency FaceFusion method that transfers identity features from the Identity Encoder to the denoising UNet of a pre-trained SD model; 3) An Attribute Rigger that injects those conditions into the denoising UNet. Our model achieves comparable or even superior performance in both identity preservation and photorealism compared to existing face editing models.
title Towards Consistent and Controllable Image Synthesis for Face Editing
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2502.02465