Beyond Words: Exploring Cultural Value Sensitivity in Multimodal Models

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Yadav, Srishti, Zhang, Zhi, Hershcovich, Daniel, Shutova, Ekaterina
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866913701101043712
author Yadav, Srishti
Zhang, Zhi
Hershcovich, Daniel
Shutova, Ekaterina
author_facet Yadav, Srishti
Zhang, Zhi
Hershcovich, Daniel
Shutova, Ekaterina
contents Investigating value alignment in Large Language Models (LLMs) based on cultural context has become a critical area of research. However, similar biases have not been extensively explored in large vision-language models (VLMs). As the scale of multimodal models continues to grow, it becomes increasingly important to assess whether images can serve as reliable proxies for culture and how these values are embedded through the integration of both visual and textual data. In this paper, we conduct a thorough evaluation of multimodal model at different scales, focusing on their alignment with cultural values. Our findings reveal that, much like LLMs, VLMs exhibit sensitivity to cultural values, but their performance in aligning with these values is highly context-dependent. While VLMs show potential in improving value understanding through the use of images, this alignment varies significantly across contexts highlighting the complexities and underexplored challenges in the alignment of multimodal models.
format Preprint
id arxiv_https___arxiv_org_abs_2502_14906
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Beyond Words: Exploring Cultural Value Sensitivity in Multimodal Models
Yadav, Srishti
Zhang, Zhi
Hershcovich, Daniel
Shutova, Ekaterina
Computation and Language
Artificial Intelligence
Investigating value alignment in Large Language Models (LLMs) based on cultural context has become a critical area of research. However, similar biases have not been extensively explored in large vision-language models (VLMs). As the scale of multimodal models continues to grow, it becomes increasingly important to assess whether images can serve as reliable proxies for culture and how these values are embedded through the integration of both visual and textual data. In this paper, we conduct a thorough evaluation of multimodal model at different scales, focusing on their alignment with cultural values. Our findings reveal that, much like LLMs, VLMs exhibit sensitivity to cultural values, but their performance in aligning with these values is highly context-dependent. While VLMs show potential in improving value understanding through the use of images, this alignment varies significantly across contexts highlighting the complexities and underexplored challenges in the alignment of multimodal models.
title Beyond Words: Exploring Cultural Value Sensitivity in Multimodal Models
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2502.14906