Visual and Textual Prompts in VLLMs for Enhancing Emotion Recognition

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Zhifeng, Zhang, Qixuan, Zhang, Peter, Niu, Wenjia, Zhang, Kaihao, Sankaranarayana, Ramesh, Caldwell, Sabrina, Gedeon, Tom
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909683665600512
author Wang, Zhifeng
Zhang, Qixuan
Zhang, Peter
Niu, Wenjia
Zhang, Kaihao
Sankaranarayana, Ramesh
Caldwell, Sabrina
Gedeon, Tom
author_facet Wang, Zhifeng
Zhang, Qixuan
Zhang, Peter
Niu, Wenjia
Zhang, Kaihao
Sankaranarayana, Ramesh
Caldwell, Sabrina
Gedeon, Tom
contents Vision Large Language Models (VLLMs) exhibit promising potential for multi-modal understanding, yet their application to video-based emotion recognition remains limited by insufficient spatial and contextual awareness. Traditional approaches, which prioritize isolated facial features, often neglect critical non-verbal cues such as body language, environmental context, and social interactions, leading to reduced robustness in real-world scenarios. To address this gap, we propose Set-of-Vision-Text Prompting (SoVTP), a novel framework that enhances zero-shot emotion recognition by integrating spatial annotations (e.g., bounding boxes, facial landmarks), physiological signals (facial action units), and contextual cues (body posture, scene dynamics, others' emotions) into a unified prompting strategy. SoVTP preserves holistic scene information while enabling fine-grained analysis of facial muscle movements and interpersonal dynamics. Extensive experiments show that SoVTP achieves substantial improvements over existing visual prompting methods, demonstrating its effectiveness in enhancing VLLMs' video emotion recognition capabilities.
format Preprint
id arxiv_https___arxiv_org_abs_2504_17224
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Visual and Textual Prompts in VLLMs for Enhancing Emotion Recognition
Wang, Zhifeng
Zhang, Qixuan
Zhang, Peter
Niu, Wenjia
Zhang, Kaihao
Sankaranarayana, Ramesh
Caldwell, Sabrina
Gedeon, Tom
Computer Vision and Pattern Recognition
Vision Large Language Models (VLLMs) exhibit promising potential for multi-modal understanding, yet their application to video-based emotion recognition remains limited by insufficient spatial and contextual awareness. Traditional approaches, which prioritize isolated facial features, often neglect critical non-verbal cues such as body language, environmental context, and social interactions, leading to reduced robustness in real-world scenarios. To address this gap, we propose Set-of-Vision-Text Prompting (SoVTP), a novel framework that enhances zero-shot emotion recognition by integrating spatial annotations (e.g., bounding boxes, facial landmarks), physiological signals (facial action units), and contextual cues (body posture, scene dynamics, others' emotions) into a unified prompting strategy. SoVTP preserves holistic scene information while enabling fine-grained analysis of facial muscle movements and interpersonal dynamics. Extensive experiments show that SoVTP achieves substantial improvements over existing visual prompting methods, demonstrating its effectiveness in enhancing VLLMs' video emotion recognition capabilities.
title Visual and Textual Prompts in VLLMs for Enhancing Emotion Recognition
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2504.17224