Robust Harmful Meme Detection under Missing Modalities via Shared Representation Learning
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866910007696556032 |
|---|---|
| author | Breiteneder, Felix Belal, Mohammad Saeed, Muhammad Saad Masoudian, Shahed Naseem, Usman Juhi, Kulshrestha Schedl, Markus Nawaz, Shah |
| author_facet | Breiteneder, Felix Belal, Mohammad Saeed, Muhammad Saad Masoudian, Shahed Naseem, Usman Juhi, Kulshrestha Schedl, Markus Nawaz, Shah |
| contents | Internet memes are powerful tools for communication, capable of spreading political, psychological, and sociocultural ideas. However, they can be harmful and can be used to disseminate hate toward targeted individuals or groups. Although previous studies have focused on designing new detection methods, these often rely on modal-complete data, such as text and images. In real-world settings, however, modalities like text may be missing due to issues like poor OCR quality, making existing methods sensitive to missing information and leading to performance deterioration. To address this gap, in this paper, we present the first-of-its-kind work to comprehensively investigate the behavior of harmful meme detection methods in the presence of modal-incomplete data. Specifically, we propose a new baseline method that learns a shared representation for multiple modalities by projecting them independently. These shared representations can then be leveraged when data is modal-incomplete. Experimental results on two benchmark datasets demonstrate that our method outperforms existing approaches when text is missing. Moreover, these results suggest that our method allows for better integration of visual features, reducing dependence on text and improving robustness in scenarios where textual information is missing. Our work represents a significant step forward in enabling the real-world application of harmful meme detection, particularly in situations where a modality is absent. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2602_01101 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Robust Harmful Meme Detection under Missing Modalities via Shared Representation Learning Breiteneder, Felix Belal, Mohammad Saeed, Muhammad Saad Masoudian, Shahed Naseem, Usman Juhi, Kulshrestha Schedl, Markus Nawaz, Shah Computer Vision and Pattern Recognition Internet memes are powerful tools for communication, capable of spreading political, psychological, and sociocultural ideas. However, they can be harmful and can be used to disseminate hate toward targeted individuals or groups. Although previous studies have focused on designing new detection methods, these often rely on modal-complete data, such as text and images. In real-world settings, however, modalities like text may be missing due to issues like poor OCR quality, making existing methods sensitive to missing information and leading to performance deterioration. To address this gap, in this paper, we present the first-of-its-kind work to comprehensively investigate the behavior of harmful meme detection methods in the presence of modal-incomplete data. Specifically, we propose a new baseline method that learns a shared representation for multiple modalities by projecting them independently. These shared representations can then be leveraged when data is modal-incomplete. Experimental results on two benchmark datasets demonstrate that our method outperforms existing approaches when text is missing. Moreover, these results suggest that our method allows for better integration of visual features, reducing dependence on text and improving robustness in scenarios where textual information is missing. Our work represents a significant step forward in enabling the real-world application of harmful meme detection, particularly in situations where a modality is absent. |
| title | Robust Harmful Meme Detection under Missing Modalities via Shared Representation Learning |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2602.01101 |