ExpertGen: Training-Free Expert Guidance for Controllable Text-to-Face Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shi, Liang, Fu, Yun
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908376156340224
author Shi, Liang
Fu, Yun
author_facet Shi, Liang
Fu, Yun
contents Recent advances in diffusion models have significantly improved text-to-face generation, but achieving fine-grained control over facial features remains a challenge. Existing methods often require training additional modules to handle specific controls such as identity, attributes, or age, making them inflexible and resource-intensive. We propose ExpertGen, a training-free framework that leverages pre-trained expert models such as face recognition, facial attribute recognition, and age estimation networks to guide generation with fine control. Our approach uses a latent consistency model to ensure realistic and in-distribution predictions at each diffusion step, enabling accurate guidance signals to effectively steer the diffusion process. We show qualitatively and quantitatively that expert models can guide the generation process with high precision, and multiple experts can collaborate to enable simultaneous control over diverse facial aspects. By allowing direct integration of off-the-shelf expert models, our method transforms any such model into a plug-and-play component for controllable face generation.
format Preprint
id arxiv_https___arxiv_org_abs_2505_17256
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ExpertGen: Training-Free Expert Guidance for Controllable Text-to-Face Generation
Shi, Liang
Fu, Yun
Computer Vision and Pattern Recognition
Recent advances in diffusion models have significantly improved text-to-face generation, but achieving fine-grained control over facial features remains a challenge. Existing methods often require training additional modules to handle specific controls such as identity, attributes, or age, making them inflexible and resource-intensive. We propose ExpertGen, a training-free framework that leverages pre-trained expert models such as face recognition, facial attribute recognition, and age estimation networks to guide generation with fine control. Our approach uses a latent consistency model to ensure realistic and in-distribution predictions at each diffusion step, enabling accurate guidance signals to effectively steer the diffusion process. We show qualitatively and quantitatively that expert models can guide the generation process with high precision, and multiple experts can collaborate to enable simultaneous control over diverse facial aspects. By allowing direct integration of off-the-shelf expert models, our method transforms any such model into a plug-and-play component for controllable face generation.
title ExpertGen: Training-Free Expert Guidance for Controllable Text-to-Face Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.17256