BLEnD-Vis: Benchmarking Multimodal Cultural Understanding in Vision Language Models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Tan, Bryan Chen Zhengyu, Weihua, Zheng, Liu, Zhengyuan, Chen, Nancy F., Lee, Hwaran, Choo, Kenny Tsu Wei, Lee, Roy Ka-Wei
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908784584032256
author Tan, Bryan Chen Zhengyu
Weihua, Zheng
Liu, Zhengyuan
Chen, Nancy F.
Lee, Hwaran
Choo, Kenny Tsu Wei
Lee, Roy Ka-Wei
author_facet Tan, Bryan Chen Zhengyu
Weihua, Zheng
Liu, Zhengyuan
Chen, Nancy F.
Lee, Hwaran
Choo, Kenny Tsu Wei
Lee, Roy Ka-Wei
contents As vision-language models (VLMs) are deployed globally, their ability to understand culturally situated knowledge becomes essential. Yet, existing evaluations largely assess static recall or isolated visual grounding, leaving unanswered whether VLMs possess robust and transferable cultural understanding. We introduce BLEnD-Vis, a multimodal, multicultural benchmark designed to evaluate the robustness of everyday cultural knowledge in VLMs across linguistic rephrasings and visual modalities. Building on the BLEnD dataset, BLEnD-Vis constructs 313 culturally grounded question templates spanning 16 regions and generates three aligned multiple-choice formats: (i) a text-only baseline querying from Region $\rightarrow$ Entity, (ii) an inverted text-only variant (Entity $\rightarrow$ Region), and (iii) a VQA-style version of (ii) with generated images. The resulting benchmark comprises 4,916 images and over 21,000 multiple-choice questions (MCQ) instances, validated through human annotation. BLEnD-Vis reveals significant fragility in current VLM cultural knowledge; models exhibit performance drops under linguistic rephrasing. While visual cues often aid performance, low cross-modal consistency highlights the challenges of robustly integrating textual and visual understanding, particularly in lower-resource regions. BLEnD-Vis thus provides a crucial testbed for systematically analysing cultural robustness and multimodal grounding, exposing limitations and guiding the development of more culturally competent VLMs. Code is available at https://github.com/Social-AI-Studio/BLEnD-Vis.
format Preprint
id arxiv_https___arxiv_org_abs_2510_11178
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle BLEnD-Vis: Benchmarking Multimodal Cultural Understanding in Vision Language Models
Tan, Bryan Chen Zhengyu
Weihua, Zheng
Liu, Zhengyuan
Chen, Nancy F.
Lee, Hwaran
Choo, Kenny Tsu Wei
Lee, Roy Ka-Wei
Computer Vision and Pattern Recognition
Computers and Society
As vision-language models (VLMs) are deployed globally, their ability to understand culturally situated knowledge becomes essential. Yet, existing evaluations largely assess static recall or isolated visual grounding, leaving unanswered whether VLMs possess robust and transferable cultural understanding. We introduce BLEnD-Vis, a multimodal, multicultural benchmark designed to evaluate the robustness of everyday cultural knowledge in VLMs across linguistic rephrasings and visual modalities. Building on the BLEnD dataset, BLEnD-Vis constructs 313 culturally grounded question templates spanning 16 regions and generates three aligned multiple-choice formats: (i) a text-only baseline querying from Region $\rightarrow$ Entity, (ii) an inverted text-only variant (Entity $\rightarrow$ Region), and (iii) a VQA-style version of (ii) with generated images. The resulting benchmark comprises 4,916 images and over 21,000 multiple-choice questions (MCQ) instances, validated through human annotation. BLEnD-Vis reveals significant fragility in current VLM cultural knowledge; models exhibit performance drops under linguistic rephrasing. While visual cues often aid performance, low cross-modal consistency highlights the challenges of robustly integrating textual and visual understanding, particularly in lower-resource regions. BLEnD-Vis thus provides a crucial testbed for systematically analysing cultural robustness and multimodal grounding, exposing limitations and guiding the development of more culturally competent VLMs. Code is available at https://github.com/Social-AI-Studio/BLEnD-Vis.
title BLEnD-Vis: Benchmarking Multimodal Cultural Understanding in Vision Language Models
topic Computer Vision and Pattern Recognition
Computers and Society
url https://arxiv.org/abs/2510.11178