Robust Harmful Meme Detection under Missing Modalities via Shared Representation Learning

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Breiteneder, Felix, Belal, Mohammad, Saeed, Muhammad Saad, Masoudian, Shahed, Naseem, Usman, Juhi, Kulshrestha, Schedl, Markus, Nawaz, Shah
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866910007696556032
author Breiteneder, Felix
Belal, Mohammad
Saeed, Muhammad Saad
Masoudian, Shahed
Naseem, Usman
Juhi, Kulshrestha
Schedl, Markus
Nawaz, Shah
author_facet Breiteneder, Felix
Belal, Mohammad
Saeed, Muhammad Saad
Masoudian, Shahed
Naseem, Usman
Juhi, Kulshrestha
Schedl, Markus
Nawaz, Shah
contents Internet memes are powerful tools for communication, capable of spreading political, psychological, and sociocultural ideas. However, they can be harmful and can be used to disseminate hate toward targeted individuals or groups. Although previous studies have focused on designing new detection methods, these often rely on modal-complete data, such as text and images. In real-world settings, however, modalities like text may be missing due to issues like poor OCR quality, making existing methods sensitive to missing information and leading to performance deterioration. To address this gap, in this paper, we present the first-of-its-kind work to comprehensively investigate the behavior of harmful meme detection methods in the presence of modal-incomplete data. Specifically, we propose a new baseline method that learns a shared representation for multiple modalities by projecting them independently. These shared representations can then be leveraged when data is modal-incomplete. Experimental results on two benchmark datasets demonstrate that our method outperforms existing approaches when text is missing. Moreover, these results suggest that our method allows for better integration of visual features, reducing dependence on text and improving robustness in scenarios where textual information is missing. Our work represents a significant step forward in enabling the real-world application of harmful meme detection, particularly in situations where a modality is absent.
format Preprint
id arxiv_https___arxiv_org_abs_2602_01101
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Robust Harmful Meme Detection under Missing Modalities via Shared Representation Learning
Breiteneder, Felix
Belal, Mohammad
Saeed, Muhammad Saad
Masoudian, Shahed
Naseem, Usman
Juhi, Kulshrestha
Schedl, Markus
Nawaz, Shah
Computer Vision and Pattern Recognition
Internet memes are powerful tools for communication, capable of spreading political, psychological, and sociocultural ideas. However, they can be harmful and can be used to disseminate hate toward targeted individuals or groups. Although previous studies have focused on designing new detection methods, these often rely on modal-complete data, such as text and images. In real-world settings, however, modalities like text may be missing due to issues like poor OCR quality, making existing methods sensitive to missing information and leading to performance deterioration. To address this gap, in this paper, we present the first-of-its-kind work to comprehensively investigate the behavior of harmful meme detection methods in the presence of modal-incomplete data. Specifically, we propose a new baseline method that learns a shared representation for multiple modalities by projecting them independently. These shared representations can then be leveraged when data is modal-incomplete. Experimental results on two benchmark datasets demonstrate that our method outperforms existing approaches when text is missing. Moreover, these results suggest that our method allows for better integration of visual features, reducing dependence on text and improving robustness in scenarios where textual information is missing. Our work represents a significant step forward in enabling the real-world application of harmful meme detection, particularly in situations where a modality is absent.
title Robust Harmful Meme Detection under Missing Modalities via Shared Representation Learning
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2602.01101