AICA-Bench: Holistically Examining the Capabilities of VLMs in Affective Image Content Analysis

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: She, Dong, Yao, Xianrong, Chen, Liqun, Yu, Jinghe, Gao, Yang, Jin, Zhanpeng
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914453740584960
author She, Dong
Yao, Xianrong
Chen, Liqun
Yu, Jinghe
Gao, Yang
Jin, Zhanpeng
author_facet She, Dong
Yao, Xianrong
Chen, Liqun
Yu, Jinghe
Gao, Yang
Jin, Zhanpeng
contents Vision-Language Models (VLMs) have demonstrated strong capabilities in perception, yet holistic Affective Image Content Analysis (AICA), which integrates perception, reasoning, and generation into a unified framework, remains underexplored. To address this gap, we introduce AICA-Bench, a comprehensive benchmark with three core tasks: Emotion Understanding (EU), Emotion Reasoning (ER), and Emotion-Guided Content Generation (EGCG). We evaluate 23 VLMs and identify two major limitations: weak intensity calibration and shallow open-ended descriptions. To address these issues, we propose Grounded Affective Tree (GAT) Prompting, a training-free framework that combines visual scaffolding with hierarchical reasoning. Experiments show that GAT reduces intensity errors and improves descriptive depth, providing a strong baseline for future research on affective multimodal understanding and generation.
format Preprint
id arxiv_https___arxiv_org_abs_2604_05900
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle AICA-Bench: Holistically Examining the Capabilities of VLMs in Affective Image Content Analysis
She, Dong
Yao, Xianrong
Chen, Liqun
Yu, Jinghe
Gao, Yang
Jin, Zhanpeng
Computer Vision and Pattern Recognition
Vision-Language Models (VLMs) have demonstrated strong capabilities in perception, yet holistic Affective Image Content Analysis (AICA), which integrates perception, reasoning, and generation into a unified framework, remains underexplored. To address this gap, we introduce AICA-Bench, a comprehensive benchmark with three core tasks: Emotion Understanding (EU), Emotion Reasoning (ER), and Emotion-Guided Content Generation (EGCG). We evaluate 23 VLMs and identify two major limitations: weak intensity calibration and shallow open-ended descriptions. To address these issues, we propose Grounded Affective Tree (GAT) Prompting, a training-free framework that combines visual scaffolding with hierarchical reasoning. Experiments show that GAT reduces intensity errors and improves descriptive depth, providing a strong baseline for future research on affective multimodal understanding and generation.
title AICA-Bench: Holistically Examining the Capabilities of VLMs in Affective Image Content Analysis
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2604.05900