FBHM: Functional Benchmarking and Steering of VLMs for Hateful Meme Detection

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Bhaskar, Paramananda, Rizwan, Naquee, Jogchand, Daksh, Pandey, Saurabh Kumar, Mukherjee, Animesh
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866914617707462656
author Bhaskar, Paramananda
Rizwan, Naquee
Jogchand, Daksh
Pandey, Saurabh Kumar
Mukherjee, Animesh
author_facet Bhaskar, Paramananda
Rizwan, Naquee
Jogchand, Daksh
Pandey, Saurabh Kumar
Mukherjee, Animesh
contents Hateful meme detection remains a formidable challenge for vision-language models, as existing benchmarks are structurally observational - confounding rhetorical hate mechanisms with target community features and preventing causal evaluation of model vulnerabilities. To address this, we introduce FBHM, a systematically curated benchmark of Functionality Based Hateful Memes constructed along two orthogonal axes: 25 distinct rhetorical functionalities and 10 target communities (5,000 memes total). Benchmarking state-of-the-art VLMs reveals a severe generalization gap: models highly accurate on standard datasets catastrophically drop to near-random performance on FBHM, proving they exploit dataset-specific heuristics rather than robust multimodal reasoning. To efficiently close this gap, we propose LSV (learnable steering vectors), an ultra-low data regime strategy that applies a causal intervention objective on as few as 500 steering samples (50 unique base memes), boosting FBHM performance by ~30 Macro-F1 points while outperforming in-context learning and PEFT without degrading source-domain performance.
format Preprint
id arxiv_https___arxiv_org_abs_2605_31349
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle FBHM: Functional Benchmarking and Steering of VLMs for Hateful Meme Detection
Bhaskar, Paramananda
Rizwan, Naquee
Jogchand, Daksh
Pandey, Saurabh Kumar
Mukherjee, Animesh
Computation and Language
Artificial Intelligence
Computer Vision and Pattern Recognition
Multimedia
Hateful meme detection remains a formidable challenge for vision-language models, as existing benchmarks are structurally observational - confounding rhetorical hate mechanisms with target community features and preventing causal evaluation of model vulnerabilities. To address this, we introduce FBHM, a systematically curated benchmark of Functionality Based Hateful Memes constructed along two orthogonal axes: 25 distinct rhetorical functionalities and 10 target communities (5,000 memes total). Benchmarking state-of-the-art VLMs reveals a severe generalization gap: models highly accurate on standard datasets catastrophically drop to near-random performance on FBHM, proving they exploit dataset-specific heuristics rather than robust multimodal reasoning. To efficiently close this gap, we propose LSV (learnable steering vectors), an ultra-low data regime strategy that applies a causal intervention objective on as few as 500 steering samples (50 unique base memes), boosting FBHM performance by ~30 Macro-F1 points while outperforming in-context learning and PEFT without degrading source-domain performance.
title FBHM: Functional Benchmarking and Steering of VLMs for Hateful Meme Detection
topic Computation and Language
Artificial Intelligence
Computer Vision and Pattern Recognition
Multimedia
url https://arxiv.org/abs/2605.31349