An Evaluation of a Visual Question Answering Strategy for Zero-shot Facial Expression Recognition in Still Images
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866908343825596416 |
|---|---|
| author | Castrillón-Santana, Modesto Santana, Oliverio J Freire-Obregón, David Hernández-Sosa, Daniel Lorenzo-Navarro, Javier |
| author_facet | Castrillón-Santana, Modesto Santana, Oliverio J Freire-Obregón, David Hernández-Sosa, Daniel Lorenzo-Navarro, Javier |
| contents | Facial expression recognition (FER) is a key research area in computer vision and human-computer interaction. Despite recent advances in deep learning, challenges persist, especially in generalizing to new scenarios. In fact, zero-shot FER significantly reduces the performance of state-of-the-art FER models. To address this problem, the community has recently started to explore the integration of knowledge from Large Language Models for visual tasks. In this work, we evaluate a broad collection of locally executed Visual Language Models (VLMs), avoiding the lack of task-specific knowledge by adopting a Visual Question Answering strategy. We compare the proposed pipeline with state-of-the-art FER models, both integrating and excluding VLMs, evaluating well-known FER benchmarks: AffectNet, FERPlus, and RAF-DB. The results show excellent performance for some VLMs in zero-shot FER scenarios, indicating the need for further exploration to improve FER generalization. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2504_21309 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | An Evaluation of a Visual Question Answering Strategy for Zero-shot Facial Expression Recognition in Still Images Castrillón-Santana, Modesto Santana, Oliverio J Freire-Obregón, David Hernández-Sosa, Daniel Lorenzo-Navarro, Javier Computer Vision and Pattern Recognition I.2.10 Facial expression recognition (FER) is a key research area in computer vision and human-computer interaction. Despite recent advances in deep learning, challenges persist, especially in generalizing to new scenarios. In fact, zero-shot FER significantly reduces the performance of state-of-the-art FER models. To address this problem, the community has recently started to explore the integration of knowledge from Large Language Models for visual tasks. In this work, we evaluate a broad collection of locally executed Visual Language Models (VLMs), avoiding the lack of task-specific knowledge by adopting a Visual Question Answering strategy. We compare the proposed pipeline with state-of-the-art FER models, both integrating and excluding VLMs, evaluating well-known FER benchmarks: AffectNet, FERPlus, and RAF-DB. The results show excellent performance for some VLMs in zero-shot FER scenarios, indicating the need for further exploration to improve FER generalization. |
| title | An Evaluation of a Visual Question Answering Strategy for Zero-shot Facial Expression Recognition in Still Images |
| topic | Computer Vision and Pattern Recognition I.2.10 |
| url | https://arxiv.org/abs/2504.21309 |