ValueGround: Evaluating Culture-Conditioned Visual Value Grounding in MLLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Zhipin, Leiter, Christoph, Frey, Christian, Abdalla, Mohamed Hesham Ibrahim, Grabocka, Josif, Eger, Steffen
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911732948008960
author Wang, Zhipin
Leiter, Christoph
Frey, Christian
Abdalla, Mohamed Hesham Ibrahim
Grabocka, Josif
Eger, Steffen
author_facet Wang, Zhipin
Leiter, Christoph
Frey, Christian
Abdalla, Mohamed Hesham Ibrahim
Grabocka, Josif
Eger, Steffen
contents Cultural values are expressed not only through language but also through visual scenes and everyday social practices. Yet existing evaluations of cultural values in language models are almost entirely text-only, leaving it unclear whether culture-conditioned judgments remain stable when response options are visualized. We introduce ValueGround, a benchmark for evaluating culture-conditioned visual value grounding in multimodal large language models (MLLMs). Built from World Values Survey questions, ValueGround uses minimally contrastive image pairs to represent opposing response options while controlling irrelevant variation. Given a country, a question, and an image pair, a model must choose the image that best matches the country's value tendency without access to the original response-option texts. Experiments across six MLLMs and 13 countries show that models perform substantially worse with visualized response options than with the original textual options, with average accuracy dropping from 72.8% to 62.6%. Our benchmark provides a controlled testbed for studying cross-modal transfer of culture-conditioned value judgments.
format Preprint
id arxiv_https___arxiv_org_abs_2604_06484
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle ValueGround: Evaluating Culture-Conditioned Visual Value Grounding in MLLMs
Wang, Zhipin
Leiter, Christoph
Frey, Christian
Abdalla, Mohamed Hesham Ibrahim
Grabocka, Josif
Eger, Steffen
Computation and Language
Cultural values are expressed not only through language but also through visual scenes and everyday social practices. Yet existing evaluations of cultural values in language models are almost entirely text-only, leaving it unclear whether culture-conditioned judgments remain stable when response options are visualized. We introduce ValueGround, a benchmark for evaluating culture-conditioned visual value grounding in multimodal large language models (MLLMs). Built from World Values Survey questions, ValueGround uses minimally contrastive image pairs to represent opposing response options while controlling irrelevant variation. Given a country, a question, and an image pair, a model must choose the image that best matches the country's value tendency without access to the original response-option texts. Experiments across six MLLMs and 13 countries show that models perform substantially worse with visualized response options than with the original textual options, with average accuracy dropping from 72.8% to 62.6%. Our benchmark provides a controlled testbed for studying cross-modal transfer of culture-conditioned value judgments.
title ValueGround: Evaluating Culture-Conditioned Visual Value Grounding in MLLMs
topic Computation and Language
url https://arxiv.org/abs/2604.06484