Mixture of Global and Local Experts with Diffusion Transformer for Controllable Face Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zou, Xuechao, Zhang, Shun, Fu, Xing, Li, Yue, Li, Kai, Cao, Yushe, Lang, Congyan, Tao, Pin, Xing, Junliang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912560265035776
author Zou, Xuechao
Zhang, Shun
Fu, Xing
Li, Yue
Li, Kai
Cao, Yushe
Lang, Congyan
Tao, Pin
Xing, Junliang
author_facet Zou, Xuechao
Zhang, Shun
Fu, Xing
Li, Yue
Li, Kai
Cao, Yushe
Lang, Congyan
Tao, Pin
Xing, Junliang
contents Controllable face generation poses critical challenges in generative modeling due to the intricate balance required between semantic controllability and photorealism. While existing approaches struggle with disentangling semantic controls from generation pipelines, we revisit the architectural potential of Diffusion Transformers (DiTs) through the lens of expert specialization. This paper introduces Face-MoGLE, a novel framework featuring: (1) Semantic-decoupled latent modeling through mask-conditioned space factorization, enabling precise attribute manipulation; (2) A mixture of global and local experts that captures holistic structure and region-level semantics for fine-grained controllability; (3) A dynamic gating network producing time-dependent coefficients that evolve with diffusion steps and spatial locations. Face-MoGLE provides a powerful and flexible solution for high-quality, controllable face generation, with strong potential in generative modeling and security applications. Extensive experiments demonstrate its effectiveness in multimodal and monomodal face generation settings and its robust zero-shot generalization capability. Project page is available at https://github.com/XavierJiezou/Face-MoGLE.
format Preprint
id arxiv_https___arxiv_org_abs_2509_00428
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Mixture of Global and Local Experts with Diffusion Transformer for Controllable Face Generation
Zou, Xuechao
Zhang, Shun
Fu, Xing
Li, Yue
Li, Kai
Cao, Yushe
Lang, Congyan
Tao, Pin
Xing, Junliang
Computer Vision and Pattern Recognition
Controllable face generation poses critical challenges in generative modeling due to the intricate balance required between semantic controllability and photorealism. While existing approaches struggle with disentangling semantic controls from generation pipelines, we revisit the architectural potential of Diffusion Transformers (DiTs) through the lens of expert specialization. This paper introduces Face-MoGLE, a novel framework featuring: (1) Semantic-decoupled latent modeling through mask-conditioned space factorization, enabling precise attribute manipulation; (2) A mixture of global and local experts that captures holistic structure and region-level semantics for fine-grained controllability; (3) A dynamic gating network producing time-dependent coefficients that evolve with diffusion steps and spatial locations. Face-MoGLE provides a powerful and flexible solution for high-quality, controllable face generation, with strong potential in generative modeling and security applications. Extensive experiments demonstrate its effectiveness in multimodal and monomodal face generation settings and its robust zero-shot generalization capability. Project page is available at https://github.com/XavierJiezou/Face-MoGLE.
title Mixture of Global and Local Experts with Diffusion Transformer for Controllable Face Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2509.00428