Are Vision Language Models Cross-Cultural Theory of Mind Reasoners?

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Nazi, Zabir Al, Shahariar, GM, Hossain, Md. Abrar, Peng, Wei
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866917187561717760
author Nazi, Zabir Al
Shahariar, GM
Hossain, Md. Abrar
Peng, Wei
author_facet Nazi, Zabir Al
Shahariar, GM
Hossain, Md. Abrar
Peng, Wei
contents Theory of Mind (ToM) - the ability to attribute beliefs and intents to others - is fundamental for social intelligence, yet Vision-Language Model (VLM) evaluations remain largely Western-centric. In this work, we introduce CulturalToM-VQA, a benchmark of 5,095 visually situated ToM probes across diverse cultural contexts, rituals, and social norms. Constructed through a frontier proprietary MLLM, human-verified pipeline, the dataset spans a taxonomy of six ToM tasks and four complexity levels. We benchmark 10 VLMs (2023-2025) and observe a significant performance leap: while earlier models struggle, frontier models achieve high accuracy (>93%). However, significant limitations persist: models struggle with false belief reasoning (19-83% accuracy) and show high regional variance (20-30% gaps). Crucially, we find that SOTA models exhibit social desirability bias - systematically favoring semantically positive answer choices over negative ones. Ablation experiments reveal that some frontier models rely heavily on parametric social priors, frequently defaulting to safety-aligned predictions. Furthermore, while Chain-of-Thought prompting aids older models, it yields minimal gains for newer ones. Overall, our work provides a testbed for cross-cultural social reasoning, underscoring that despite architectural gains, achieving robust, visually grounded understanding remains an open challenge.
format Preprint
id arxiv_https___arxiv_org_abs_2512_17394
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Are Vision Language Models Cross-Cultural Theory of Mind Reasoners?
Nazi, Zabir Al
Shahariar, GM
Hossain, Md. Abrar
Peng, Wei
Computation and Language
Computer Vision and Pattern Recognition
Computers and Society
Theory of Mind (ToM) - the ability to attribute beliefs and intents to others - is fundamental for social intelligence, yet Vision-Language Model (VLM) evaluations remain largely Western-centric. In this work, we introduce CulturalToM-VQA, a benchmark of 5,095 visually situated ToM probes across diverse cultural contexts, rituals, and social norms. Constructed through a frontier proprietary MLLM, human-verified pipeline, the dataset spans a taxonomy of six ToM tasks and four complexity levels. We benchmark 10 VLMs (2023-2025) and observe a significant performance leap: while earlier models struggle, frontier models achieve high accuracy (>93%). However, significant limitations persist: models struggle with false belief reasoning (19-83% accuracy) and show high regional variance (20-30% gaps). Crucially, we find that SOTA models exhibit social desirability bias - systematically favoring semantically positive answer choices over negative ones. Ablation experiments reveal that some frontier models rely heavily on parametric social priors, frequently defaulting to safety-aligned predictions. Furthermore, while Chain-of-Thought prompting aids older models, it yields minimal gains for newer ones. Overall, our work provides a testbed for cross-cultural social reasoning, underscoring that despite architectural gains, achieving robust, visually grounded understanding remains an open challenge.
title Are Vision Language Models Cross-Cultural Theory of Mind Reasoners?
topic Computation and Language
Computer Vision and Pattern Recognition
Computers and Society
url https://arxiv.org/abs/2512.17394