Cultural Awareness in Vision-Language Models: A Cross-Country Exploration

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Madasu, Avinash, Lal, Vasudev, Howard, Phillip
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909624057200640
author Madasu, Avinash
Lal, Vasudev
Howard, Phillip
author_facet Madasu, Avinash
Lal, Vasudev
Howard, Phillip
contents Vision-Language Models (VLMs) are increasingly deployed in diverse cultural contexts, yet their internal biases remain poorly understood. In this work, we propose a novel framework to systematically evaluate how VLMs encode cultural differences and biases related to race, gender, and physical traits across countries. We introduce three retrieval-based tasks: (1) Race to Country retrieval, which examines the association between individuals from specific racial groups (East Asian, White, Middle Eastern, Latino, South Asian, and Black) and different countries; (2) Personal Traits to Country retrieval, where images are paired with trait-based prompts (e.g., Smart, Honest, Criminal, Violent) to investigate potential stereotypical associations; and (3) Physical Characteristics to Country retrieval, focusing on visual attributes like skinny, young, obese, and old to explore how physical appearances are culturally linked to nations. Our findings reveal persistent biases in VLMs, highlighting how visual representations may inadvertently reinforce societal stereotypes.
format Preprint
id arxiv_https___arxiv_org_abs_2505_20326
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Cultural Awareness in Vision-Language Models: A Cross-Country Exploration
Madasu, Avinash
Lal, Vasudev
Howard, Phillip
Computers and Society
Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
Vision-Language Models (VLMs) are increasingly deployed in diverse cultural contexts, yet their internal biases remain poorly understood. In this work, we propose a novel framework to systematically evaluate how VLMs encode cultural differences and biases related to race, gender, and physical traits across countries. We introduce three retrieval-based tasks: (1) Race to Country retrieval, which examines the association between individuals from specific racial groups (East Asian, White, Middle Eastern, Latino, South Asian, and Black) and different countries; (2) Personal Traits to Country retrieval, where images are paired with trait-based prompts (e.g., Smart, Honest, Criminal, Violent) to investigate potential stereotypical associations; and (3) Physical Characteristics to Country retrieval, focusing on visual attributes like skinny, young, obese, and old to explore how physical appearances are culturally linked to nations. Our findings reveal persistent biases in VLMs, highlighting how visual representations may inadvertently reinforce societal stereotypes.
title Cultural Awareness in Vision-Language Models: A Cross-Country Exploration
topic Computers and Society
Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.20326