UniF$^2$ace: A Unified Fine-grained Face Understanding and Generation Model
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Junzhe, Zhou, Sifan, Guo, Liya, Qiu, Xuerui, Xu, Linrui, Qu, Delin, Long, Tingting, Fan, Chun, Li, Ming, Fan, Hehe, Liu, Jun, Yan, Shuicheng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
UniPath: Adaptive Coordination of Understanding and Generation for Unified Multimodal Reasoning
by: Bai, Hayes, et al.
Published: (2026)
by: Bai, Hayes, et al.
Published: (2026)
UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation
by: Chen, Yanzhe, et al.
Published: (2025)
by: Chen, Yanzhe, et al.
Published: (2025)
UniSLU: Unified Spoken Language Understanding from Heterogeneous Cross-Task Datasets
by: Sheng, Zhichao, et al.
Published: (2025)
by: Sheng, Zhichao, et al.
Published: (2025)
UniSRCodec: Unified and Low-Bitrate Single Codebook Codec with Sub-Band Reconstruction
by: Zhang, Zhisheng, et al.
Published: (2026)
by: Zhang, Zhisheng, et al.
Published: (2026)
Fine-grained Knowledge Graph-driven Video-Language Learning for Action Recognition
by: Zhang, Rui, et al.
Published: (2024)
by: Zhang, Rui, et al.
Published: (2024)
Feedback-Driven Rate Control for Learned Video Compression
by: Xu, Zhiheng, et al.
Published: (2026)
by: Xu, Zhiheng, et al.
Published: (2026)
FineBadminton: A Multi-Level Dataset for Fine-Grained Badminton Video Understanding
by: He, Xusheng, et al.
Published: (2025)
by: He, Xusheng, et al.
Published: (2025)
TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic Videos
by: Kong, Fanheng, et al.
Published: (2025)
by: Kong, Fanheng, et al.
Published: (2025)
Towards Unified Representation of Multi-Modal Pre-training for 3D Understanding via Differentiable Rendering
by: Fei, Ben, et al.
Published: (2024)
by: Fei, Ben, et al.
Published: (2024)
Uni-Retrieval: A Multi-Style Retrieval Framework for STEM's Education
by: Jia, Yanhao, et al.
Published: (2025)
by: Jia, Yanhao, et al.
Published: (2025)
High-level Codes and Fine-grained Weights for Online Multi-modal Hashing Retrieval
by: Zhan, Yu-Wei, et al.
Published: (2024)
by: Zhan, Yu-Wei, et al.
Published: (2024)
Voices, Faces, and Feelings: Multi-modal Emotion-Cognition Captioning for Mental Health Understanding
by: Zhou, Zhiyuan, et al.
Published: (2026)
by: Zhou, Zhiyuan, et al.
Published: (2026)
DRFormer: A Dual-Regularized Bidirectional Transformer for Person Re-identification
by: Shu, Ying, et al.
Published: (2026)
by: Shu, Ying, et al.
Published: (2026)
PixCLIP: Achieving Fine-grained Visual Language Understanding via Any-granularity Pixel-Text Alignment Learning
by: Xiao, Yicheng, et al.
Published: (2025)
by: Xiao, Yicheng, et al.
Published: (2025)
SpeechCraft: A Fine-grained Expressive Speech Dataset with Natural Language Description
by: Jin, Zeyu, et al.
Published: (2024)
by: Jin, Zeyu, et al.
Published: (2024)
Fine-grained Image Retrieval via Dual-Vision Adaptation
by: Jiang, Xin, et al.
Published: (2025)
by: Jiang, Xin, et al.
Published: (2025)
UniTalking: A Unified Audio-Video Framework for Talking Portrait Generation
by: Li, Hebeizi, et al.
Published: (2026)
by: Li, Hebeizi, et al.
Published: (2026)
G4G:A Generic Framework for High Fidelity Talking Face Generation with Fine-grained Intra-modal Alignment
by: Zhang, Juan, et al.
Published: (2024)
by: Zhang, Juan, et al.
Published: (2024)
PC-JND: Subjective Study and Dataset on Just Noticeable Difference for Point Clouds in 6DoF Virtual Reality
by: Fan, Chunling, et al.
Published: (2025)
by: Fan, Chunling, et al.
Published: (2025)
Fostering Emotional Perspective-Taking: An Exploration of Affective Face-Tracking Interactions in the VR Narrative Rekindle
by: Fan, Hector, et al.
Published: (2026)
by: Fan, Hector, et al.
Published: (2026)
DS-HGCN: A Dual-Stream Hypergraph Convolutional Network for Predicting Student Engagement via Social Contagion
by: Fan, Ziyang, et al.
Published: (2025)
by: Fan, Ziyang, et al.
Published: (2025)
SFQA: A Comprehensive Perceptual Quality Assessment Dataset for Singing Face Generation
by: Gao, Zhilin, et al.
Published: (2026)
by: Gao, Zhilin, et al.
Published: (2026)
Identity-Aware Vision-Language Model for Explainable Face Forgery Detection
by: Xu, Junhao, et al.
Published: (2025)
by: Xu, Junhao, et al.
Published: (2025)
UMETTS: A Unified Framework for Emotional Text-to-Speech Synthesis with Multimodal Prompts
by: Cheng, Zhi-Qi, et al.
Published: (2024)
by: Cheng, Zhi-Qi, et al.
Published: (2024)
Emotion-Qwen: A Unified Framework for Emotion and Vision Understanding
by: Huang, Dawei, et al.
Published: (2025)
by: Huang, Dawei, et al.
Published: (2025)
SRA: Semantic Relation-Aware Flowchart Question Answering
by: Li, Xinyu, et al.
Published: (2026)
by: Li, Xinyu, et al.
Published: (2026)
UniCVR: From Alignment to Reranking for Unified Zero-Shot Composed Visual Retrieval
by: Wen, Haokun, et al.
Published: (2026)
by: Wen, Haokun, et al.
Published: (2026)
Rethinking Vision Transformer for Large-Scale Fine-Grained Image Retrieval
by: Jiang, Xin, et al.
Published: (2025)
by: Jiang, Xin, et al.
Published: (2025)
Uni-MoE: Scaling Unified Multimodal LLMs with Mixture of Experts
by: Li, Yunxin, et al.
Published: (2024)
by: Li, Yunxin, et al.
Published: (2024)
Universal Organizer of SAM for Unsupervised Semantic Segmentation
by: Li, Tingting, et al.
Published: (2024)
by: Li, Tingting, et al.
Published: (2024)
CEM-Net: Cross-Emotion Memory Network for Emotional Talking Face Generation
by: Wu, Kangyi, et al.
Published: (2025)
by: Wu, Kangyi, et al.
Published: (2025)
Fine-grained Image Quality Assessment for Perceptual Image Restoration
by: Sheng, Xiangfei, et al.
Published: (2025)
by: Sheng, Xiangfei, et al.
Published: (2025)
MolCA: Molecular Graph-Language Modeling with Cross-Modal Projector and Uni-Modal Adapter
by: Liu, Zhiyuan, et al.
Published: (2023)
by: Liu, Zhiyuan, et al.
Published: (2023)
Mixture-of-Prompt-Experts for Multi-modal Semantic Understanding
by: Wu, Zichen, et al.
Published: (2024)
by: Wu, Zichen, et al.
Published: (2024)
Multi-Reference Generative Face Video Compression with Contrastive Learning
by: Konuko, Goluck, et al.
Published: (2024)
by: Konuko, Goluck, et al.
Published: (2024)
MORE-R1: Guiding LVLM for Multimodal Object-Entity Relation Extraction via Stepwise Reasoning with Reinforcement Learning
by: Yuan, Xiang, et al.
Published: (2026)
by: Yuan, Xiang, et al.
Published: (2026)
Rethink Web Service Resilience in Space: A Radiation-Aware and Sustainable Transmission Solution
by: Chen, Long, et al.
Published: (2026)
by: Chen, Long, et al.
Published: (2026)
Audio-Visual Cross-Modal Compression for Generative Face Video Coding
by: Xu, Youmin, et al.
Published: (2025)
by: Xu, Youmin, et al.
Published: (2025)
UniPT: Universal Parallel Tuning for Transfer Learning with Efficient Parameter and Memory
by: Diao, Haiwen, et al.
Published: (2023)
by: Diao, Haiwen, et al.
Published: (2023)
HADUA: Hierarchical Attention and Dynamic Uniform Alignment for Robust Cross-Subject Emotion Recognition
by: Tang, Jiahao, et al.
Published: (2026)
by: Tang, Jiahao, et al.
Published: (2026)
Similar Items
-
UniPath: Adaptive Coordination of Understanding and Generation for Unified Multimodal Reasoning
by: Bai, Hayes, et al.
Published: (2026) -
UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation
by: Chen, Yanzhe, et al.
Published: (2025) -
UniSLU: Unified Spoken Language Understanding from Heterogeneous Cross-Task Datasets
by: Sheng, Zhichao, et al.
Published: (2025) -
UniSRCodec: Unified and Low-Bitrate Single Codebook Codec with Sub-Band Reconstruction
by: Zhang, Zhisheng, et al.
Published: (2026) -
Fine-grained Knowledge Graph-driven Video-Language Learning for Action Recognition
by: Zhang, Rui, et al.
Published: (2024)