Face-LLaVA: Facial Expression and Attribute Understanding through Instruction Tuning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chaubey, Ashutosh, Guan, Xulang, Soleymani, Mohammad
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915235639590912
author Chaubey, Ashutosh
Guan, Xulang
Soleymani, Mohammad
author_facet Chaubey, Ashutosh
Guan, Xulang
Soleymani, Mohammad
contents The human face plays a central role in social communication, necessitating the use of performant computer vision tools for human-centered applications. We propose Face-LLaVA, a multimodal large language model for face-centered, in-context learning, including facial expression and attribute recognition. Additionally, Face-LLaVA is able to generate natural language descriptions that can be used for reasoning. Leveraging existing visual databases, we first developed FaceInstruct-1M, a face-centered database for instruction tuning MLLMs for face processing. We then developed a novel face-specific visual encoder powered by Face-Region Guided Cross-Attention that integrates face geometry with local visual features. We evaluated the proposed method across nine different datasets and five different face processing tasks, including facial expression recognition, action unit detection, facial attribute detection, age estimation and deepfake detection. Face-LLaVA achieves superior results compared to existing open-source MLLMs and competitive performance compared to commercial solutions. Our model output also receives a higher reasoning rating by GPT under a zero-shot setting across all the tasks. Both our dataset and model wil be released at https://face-llava.github.io to support future advancements in social AI and foundational vision-language research.
format Preprint
id arxiv_https___arxiv_org_abs_2504_07198
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Face-LLaVA: Facial Expression and Attribute Understanding through Instruction Tuning
Chaubey, Ashutosh
Guan, Xulang
Soleymani, Mohammad
Computer Vision and Pattern Recognition
Artificial Intelligence
Human-Computer Interaction
The human face plays a central role in social communication, necessitating the use of performant computer vision tools for human-centered applications. We propose Face-LLaVA, a multimodal large language model for face-centered, in-context learning, including facial expression and attribute recognition. Additionally, Face-LLaVA is able to generate natural language descriptions that can be used for reasoning. Leveraging existing visual databases, we first developed FaceInstruct-1M, a face-centered database for instruction tuning MLLMs for face processing. We then developed a novel face-specific visual encoder powered by Face-Region Guided Cross-Attention that integrates face geometry with local visual features. We evaluated the proposed method across nine different datasets and five different face processing tasks, including facial expression recognition, action unit detection, facial attribute detection, age estimation and deepfake detection. Face-LLaVA achieves superior results compared to existing open-source MLLMs and competitive performance compared to commercial solutions. Our model output also receives a higher reasoning rating by GPT under a zero-shot setting across all the tasks. Both our dataset and model wil be released at https://face-llava.github.io to support future advancements in social AI and foundational vision-language research.
title Face-LLaVA: Facial Expression and Attribute Understanding through Instruction Tuning
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Human-Computer Interaction
url https://arxiv.org/abs/2504.07198