Evaluation of Cultural Competence of Vision-Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866911105016659968 |
|---|---|
| author | Yadav, Srishti Tilton, Lauren Antoniak, Maria Arnold, Taylor Li, Jiaang Pawar, Siddhesh Milind Karamolegkou, Antonia Frank, Stella An, Zhaochong Rostamzadeh, Negar Hershcovich, Daniel Belongie, Serge Shutova, Ekaterina |
| author_facet | Yadav, Srishti Tilton, Lauren Antoniak, Maria Arnold, Taylor Li, Jiaang Pawar, Siddhesh Milind Karamolegkou, Antonia Frank, Stella An, Zhaochong Rostamzadeh, Negar Hershcovich, Daniel Belongie, Serge Shutova, Ekaterina |
| contents | Modern vision-language models (VLMs) often fail at cultural competency evaluations and benchmarks. Given the diversity of applications built upon VLMs, there is renewed interest in understanding how they encode cultural nuances. While individual aspects of this problem have been studied, we still lack a comprehensive framework for systematically identifying and annotating the nuanced cultural dimensions present in images for VLMs. This position paper argues that foundational methodologies from visual culture studies (cultural studies, semiotics, and visual studies) are necessary for cultural analysis of images. Building upon this review, we propose a set of five frameworks, corresponding to cultural dimensions, that must be considered for a more complete analysis of the cultural competencies of VLMs. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2505_22793 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Evaluation of Cultural Competence of Vision-Language Models Yadav, Srishti Tilton, Lauren Antoniak, Maria Arnold, Taylor Li, Jiaang Pawar, Siddhesh Milind Karamolegkou, Antonia Frank, Stella An, Zhaochong Rostamzadeh, Negar Hershcovich, Daniel Belongie, Serge Shutova, Ekaterina Computer Vision and Pattern Recognition Computation and Language Modern vision-language models (VLMs) often fail at cultural competency evaluations and benchmarks. Given the diversity of applications built upon VLMs, there is renewed interest in understanding how they encode cultural nuances. While individual aspects of this problem have been studied, we still lack a comprehensive framework for systematically identifying and annotating the nuanced cultural dimensions present in images for VLMs. This position paper argues that foundational methodologies from visual culture studies (cultural studies, semiotics, and visual studies) are necessary for cultural analysis of images. Building upon this review, we propose a set of five frameworks, corresponding to cultural dimensions, that must be considered for a more complete analysis of the cultural competencies of VLMs. |
| title | Evaluation of Cultural Competence of Vision-Language Models |
| topic | Computer Vision and Pattern Recognition Computation and Language |
| url | https://arxiv.org/abs/2505.22793 |