Culture In a Frame: C$^3$B as a Comic-Based Benchmark for Multimodal Culturally Awareness

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Song, Yuchen, Chen, Andong, Zhu, Wenxin, Chen, Kehai, Bai, Xuefeng, Yang, Muyun, Zhao, Tiejun
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908856496422912
author Song, Yuchen
Chen, Andong
Zhu, Wenxin
Chen, Kehai
Bai, Xuefeng
Yang, Muyun
Zhao, Tiejun
author_facet Song, Yuchen
Chen, Andong
Zhu, Wenxin
Chen, Kehai
Bai, Xuefeng
Yang, Muyun
Zhao, Tiejun
contents Cultural awareness capabilities have emerged as a critical capability for Multimodal Large Language Models (MLLMs). However, current benchmarks lack progressed difficulty in their task design and are deficient in cross-lingual tasks. Moreover, current benchmarks often use real-world images. Each real-world image typically contains one culture, making these benchmarks relatively easy for MLLMs. Based on this, we propose C$^3$B (Comics Cross-Cultural Benchmark), a novel multicultural, multitask and multilingual cultural awareness capabilities benchmark. C$^3$B comprises over 2000 images and over 18000 QA pairs, constructed on three tasks with progressed difficulties, from basic visual recognition to higher-level cultural conflict understanding, and finally to cultural content generation. We conducted evaluations on 11 open-source MLLMs, revealing a significant performance gap between MLLMs and human performance. The gap demonstrates that C$^3$B poses substantial challenges for current MLLMs, encouraging future research to advance the cultural awareness capabilities of MLLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2510_00041
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Culture In a Frame: C$^3$B as a Comic-Based Benchmark for Multimodal Culturally Awareness
Song, Yuchen
Chen, Andong
Zhu, Wenxin
Chen, Kehai
Bai, Xuefeng
Yang, Muyun
Zhao, Tiejun
Computer Vision and Pattern Recognition
Artificial Intelligence
Cultural awareness capabilities have emerged as a critical capability for Multimodal Large Language Models (MLLMs). However, current benchmarks lack progressed difficulty in their task design and are deficient in cross-lingual tasks. Moreover, current benchmarks often use real-world images. Each real-world image typically contains one culture, making these benchmarks relatively easy for MLLMs. Based on this, we propose C$^3$B (Comics Cross-Cultural Benchmark), a novel multicultural, multitask and multilingual cultural awareness capabilities benchmark. C$^3$B comprises over 2000 images and over 18000 QA pairs, constructed on three tasks with progressed difficulties, from basic visual recognition to higher-level cultural conflict understanding, and finally to cultural content generation. We conducted evaluations on 11 open-source MLLMs, revealing a significant performance gap between MLLMs and human performance. The gap demonstrates that C$^3$B poses substantial challenges for current MLLMs, encouraging future research to advance the cultural awareness capabilities of MLLMs.
title Culture In a Frame: C$^3$B as a Comic-Based Benchmark for Multimodal Culturally Awareness
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2510.00041