UniF$^2$ace: A Unified Fine-grained Face Understanding and Generation Model

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Li, Junzhe, Zhou, Sifan, Guo, Liya, Qiu, Xuerui, Xu, Linrui, Qu, Delin, Long, Tingting, Fan, Chun, Li, Ming, Fan, Hehe, Liu, Jun, Yan, Shuicheng
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909988066164736
author Li, Junzhe
Zhou, Sifan
Guo, Liya
Qiu, Xuerui
Xu, Linrui
Qu, Delin
Long, Tingting
Fan, Chun
Li, Ming
Fan, Hehe
Liu, Jun
Yan, Shuicheng
author_facet Li, Junzhe
Zhou, Sifan
Guo, Liya
Qiu, Xuerui
Xu, Linrui
Qu, Delin
Long, Tingting
Fan, Chun
Li, Ming
Fan, Hehe
Liu, Jun
Yan, Shuicheng
contents Unified multimodal models (UMMs) have emerged as a powerful paradigm in fundamental cross-modality research, demonstrating significant potential in both image understanding and generation. However, existing research in the face domain primarily faces two challenges: $\textbf{(1)}$ $\textbf{fragmentation development}$, with existing methods failing to unify understanding and generation into a single one, hindering the way to artificial general intelligence. $\textbf{(2) lack of fine-grained facial attributes}$, which are crucial for high-fidelity applications. To handle those issues, we propose $\textbf{UniF$^2$ace}$, $\textit{the first UMM specifically tailored for fine-grained face understanding and generation}$. $\textbf{First}$, we introduce a novel theoretical framework with a Dual Discrete Diffusion (D3Diff) loss, unifying masked generative models with discrete score matching diffusion and leading to a more precise approximation of the negative log-likelihood. Moreover, this D3Diff significantly enhances the model's ability to synthesize high-fidelity facial details aligned with text input. $\textbf{Second}$, we propose a multi-level grouped Mixture-of-Experts architecture, adaptively incorporating the semantic and identity facial embeddings to complement the attribute forgotten phenomenon in representation evolvement. $\textbf{Finally}$, to this end, we construct UniF$^2$aceD-1M, a large-scale dataset comprising 130K fine-grained image-caption pairs and 1M visual question-answering pairs, spanning a much wider range of facial attributes than existing datasets. Extensive experiments demonstrate that UniF$^2$ace outperforms existing models with a similar scale in both understanding and generation tasks, with 7.1\% higher Desc-GPT and 6.6\% higher VQA-score, respectively.
format Preprint
id arxiv_https___arxiv_org_abs_2503_08120
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle UniF$^2$ace: A Unified Fine-grained Face Understanding and Generation Model
Li, Junzhe
Zhou, Sifan
Guo, Liya
Qiu, Xuerui
Xu, Linrui
Qu, Delin
Long, Tingting
Fan, Chun
Li, Ming
Fan, Hehe
Liu, Jun
Yan, Shuicheng
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Multimedia
Unified multimodal models (UMMs) have emerged as a powerful paradigm in fundamental cross-modality research, demonstrating significant potential in both image understanding and generation. However, existing research in the face domain primarily faces two challenges: $\textbf{(1)}$ $\textbf{fragmentation development}$, with existing methods failing to unify understanding and generation into a single one, hindering the way to artificial general intelligence. $\textbf{(2) lack of fine-grained facial attributes}$, which are crucial for high-fidelity applications. To handle those issues, we propose $\textbf{UniF$^2$ace}$, $\textit{the first UMM specifically tailored for fine-grained face understanding and generation}$. $\textbf{First}$, we introduce a novel theoretical framework with a Dual Discrete Diffusion (D3Diff) loss, unifying masked generative models with discrete score matching diffusion and leading to a more precise approximation of the negative log-likelihood. Moreover, this D3Diff significantly enhances the model's ability to synthesize high-fidelity facial details aligned with text input. $\textbf{Second}$, we propose a multi-level grouped Mixture-of-Experts architecture, adaptively incorporating the semantic and identity facial embeddings to complement the attribute forgotten phenomenon in representation evolvement. $\textbf{Finally}$, to this end, we construct UniF$^2$aceD-1M, a large-scale dataset comprising 130K fine-grained image-caption pairs and 1M visual question-answering pairs, spanning a much wider range of facial attributes than existing datasets. Extensive experiments demonstrate that UniF$^2$ace outperforms existing models with a similar scale in both understanding and generation tasks, with 7.1\% higher Desc-GPT and 6.6\% higher VQA-score, respectively.
title UniF$^2$ace: A Unified Fine-grained Face Understanding and Generation Model
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Multimedia
url https://arxiv.org/abs/2503.08120