From Native Memes to Global Moderation: Cross-Cultural Evaluation of Vision-Language Models for Hateful Meme Detection

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wang, Mo, Ren, Kaixuan, Jalan, Pratik, Ashraf, Ahmed, Vu, Tuong Vy, Seetharaman, Rahul, Nawaz, Shah, Naseem, Usman
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911442087706624
author Wang, Mo
Ren, Kaixuan
Jalan, Pratik
Ashraf, Ahmed
Vu, Tuong Vy
Seetharaman, Rahul
Nawaz, Shah
Naseem, Usman
author_facet Wang, Mo
Ren, Kaixuan
Jalan, Pratik
Ashraf, Ahmed
Vu, Tuong Vy
Seetharaman, Rahul
Nawaz, Shah
Naseem, Usman
contents Cultural context profoundly shapes how people interpret online content, yet vision-language models (VLMs) remain predominantly trained through Western or English-centric lenses. This limits their fairness and cross-cultural robustness in tasks like hateful meme detection. We introduce a systematic evaluation framework designed to diagnose and quantify the cross-cultural robustness of state-of-the-art VLMs across multilingual meme datasets, analyzing three axes: (i) learning strategy (zero-shot vs. one-shot), (ii) prompting language (native vs. English), and (iii) translation effects on meaning and detection. Results show that the common ``translate-then-detect'' approach deteriorate performance, while culturally aligned interventions - native-language prompting and one-shot learning - significantly enhance detection. Our findings reveal systematic convergence toward Western safety norms and provide actionable strategies to mitigate such bias, guiding the design of globally robust multimodal moderation systems.
format Preprint
id arxiv_https___arxiv_org_abs_2602_07497
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle From Native Memes to Global Moderation: Cross-Cultural Evaluation of Vision-Language Models for Hateful Meme Detection
Wang, Mo
Ren, Kaixuan
Jalan, Pratik
Ashraf, Ahmed
Vu, Tuong Vy
Seetharaman, Rahul
Nawaz, Shah
Naseem, Usman
Computation and Language
H.3.3; I.2.7; K.4.1
Cultural context profoundly shapes how people interpret online content, yet vision-language models (VLMs) remain predominantly trained through Western or English-centric lenses. This limits their fairness and cross-cultural robustness in tasks like hateful meme detection. We introduce a systematic evaluation framework designed to diagnose and quantify the cross-cultural robustness of state-of-the-art VLMs across multilingual meme datasets, analyzing three axes: (i) learning strategy (zero-shot vs. one-shot), (ii) prompting language (native vs. English), and (iii) translation effects on meaning and detection. Results show that the common ``translate-then-detect'' approach deteriorate performance, while culturally aligned interventions - native-language prompting and one-shot learning - significantly enhance detection. Our findings reveal systematic convergence toward Western safety norms and provide actionable strategies to mitigate such bias, guiding the design of globally robust multimodal moderation systems.
title From Native Memes to Global Moderation: Cross-Cultural Evaluation of Vision-Language Models for Hateful Meme Detection
topic Computation and Language
H.3.3; I.2.7; K.4.1
url https://arxiv.org/abs/2602.07497