An Evaluation of a Visual Question Answering Strategy for Zero-shot Facial Expression Recognition in Still Images

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Castrillón-Santana, Modesto, Santana, Oliverio J, Freire-Obregón, David, Hernández-Sosa, Daniel, Lorenzo-Navarro, Javier
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908343825596416
author Castrillón-Santana, Modesto
Santana, Oliverio J
Freire-Obregón, David
Hernández-Sosa, Daniel
Lorenzo-Navarro, Javier
author_facet Castrillón-Santana, Modesto
Santana, Oliverio J
Freire-Obregón, David
Hernández-Sosa, Daniel
Lorenzo-Navarro, Javier
contents Facial expression recognition (FER) is a key research area in computer vision and human-computer interaction. Despite recent advances in deep learning, challenges persist, especially in generalizing to new scenarios. In fact, zero-shot FER significantly reduces the performance of state-of-the-art FER models. To address this problem, the community has recently started to explore the integration of knowledge from Large Language Models for visual tasks. In this work, we evaluate a broad collection of locally executed Visual Language Models (VLMs), avoiding the lack of task-specific knowledge by adopting a Visual Question Answering strategy. We compare the proposed pipeline with state-of-the-art FER models, both integrating and excluding VLMs, evaluating well-known FER benchmarks: AffectNet, FERPlus, and RAF-DB. The results show excellent performance for some VLMs in zero-shot FER scenarios, indicating the need for further exploration to improve FER generalization.
format Preprint
id arxiv_https___arxiv_org_abs_2504_21309
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle An Evaluation of a Visual Question Answering Strategy for Zero-shot Facial Expression Recognition in Still Images
Castrillón-Santana, Modesto
Santana, Oliverio J
Freire-Obregón, David
Hernández-Sosa, Daniel
Lorenzo-Navarro, Javier
Computer Vision and Pattern Recognition
I.2.10
Facial expression recognition (FER) is a key research area in computer vision and human-computer interaction. Despite recent advances in deep learning, challenges persist, especially in generalizing to new scenarios. In fact, zero-shot FER significantly reduces the performance of state-of-the-art FER models. To address this problem, the community has recently started to explore the integration of knowledge from Large Language Models for visual tasks. In this work, we evaluate a broad collection of locally executed Visual Language Models (VLMs), avoiding the lack of task-specific knowledge by adopting a Visual Question Answering strategy. We compare the proposed pipeline with state-of-the-art FER models, both integrating and excluding VLMs, evaluating well-known FER benchmarks: AffectNet, FERPlus, and RAF-DB. The results show excellent performance for some VLMs in zero-shot FER scenarios, indicating the need for further exploration to improve FER generalization.
title An Evaluation of a Visual Question Answering Strategy for Zero-shot Facial Expression Recognition in Still Images
topic Computer Vision and Pattern Recognition
I.2.10
url https://arxiv.org/abs/2504.21309