AICA-Bench: Holistically Examining the Capabilities of VLMs in Affective Image Content Analysis
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866914453740584960 |
|---|---|
| author | She, Dong Yao, Xianrong Chen, Liqun Yu, Jinghe Gao, Yang Jin, Zhanpeng |
| author_facet | She, Dong Yao, Xianrong Chen, Liqun Yu, Jinghe Gao, Yang Jin, Zhanpeng |
| contents | Vision-Language Models (VLMs) have demonstrated strong capabilities in perception, yet holistic Affective Image Content Analysis (AICA), which integrates perception, reasoning, and generation into a unified framework, remains underexplored. To address this gap, we introduce AICA-Bench, a comprehensive benchmark with three core tasks: Emotion Understanding (EU), Emotion Reasoning (ER), and Emotion-Guided Content Generation (EGCG). We evaluate 23 VLMs and identify two major limitations: weak intensity calibration and shallow open-ended descriptions. To address these issues, we propose Grounded Affective Tree (GAT) Prompting, a training-free framework that combines visual scaffolding with hierarchical reasoning. Experiments show that GAT reduces intensity errors and improves descriptive depth, providing a strong baseline for future research on affective multimodal understanding and generation. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2604_05900 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | AICA-Bench: Holistically Examining the Capabilities of VLMs in Affective Image Content Analysis She, Dong Yao, Xianrong Chen, Liqun Yu, Jinghe Gao, Yang Jin, Zhanpeng Computer Vision and Pattern Recognition Vision-Language Models (VLMs) have demonstrated strong capabilities in perception, yet holistic Affective Image Content Analysis (AICA), which integrates perception, reasoning, and generation into a unified framework, remains underexplored. To address this gap, we introduce AICA-Bench, a comprehensive benchmark with three core tasks: Emotion Understanding (EU), Emotion Reasoning (ER), and Emotion-Guided Content Generation (EGCG). We evaluate 23 VLMs and identify two major limitations: weak intensity calibration and shallow open-ended descriptions. To address these issues, we propose Grounded Affective Tree (GAT) Prompting, a training-free framework that combines visual scaffolding with hierarchical reasoning. Experiments show that GAT reduces intensity errors and improves descriptive depth, providing a strong baseline for future research on affective multimodal understanding and generation. |
| title | AICA-Bench: Holistically Examining the Capabilities of VLMs in Affective Image Content Analysis |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2604.05900 |