Saved in:
Bibliographic Details
Main Authors: Sharshar, Ahmed, Elgendy, Hosam, Ahmed, Saad El Dine, Rohaim, Yasser, Wang, Yuxia
Format: Preprint
Published: 2026
Subjects:
Online Access:https://arxiv.org/abs/2603.17759
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910059949195264
author Sharshar, Ahmed
Elgendy, Hosam
Ahmed, Saad El Dine
Rohaim, Yasser
Wang, Yuxia
author_facet Sharshar, Ahmed
Elgendy, Hosam
Ahmed, Saad El Dine
Rohaim, Yasser
Wang, Yuxia
contents Dark humor often relies on subtle cultural nuances and implicit cues that require contextual reasoning to interpret, posing safety challenges that current static benchmarks fail to capture. To address this, we introduce a novel multimodal, multilingual benchmark for detecting and understanding harmful and offensive humor. Our manually curated dataset comprises 3,000 texts and 6,000 images in English and Arabic, alongside 1,200 videos that span English, Arabic, and language-independent (universal) contexts. Unlike standard toxicity datasets, we enforce a strict annotation guideline: distinguishing Safe jokes from Harmful ones, with the latter further classified into Explicit (overt) and Implicit (Covert) categories to probe deep reasoning. We systematically evaluate state-of-the-art (SOTA) open and closed-source models across all modalities. Our findings reveal that closed-source models significantly outperform open-source ones, with a notable difference in performance between the English and Arabic languages in both, underscoring the critical need for culturally grounded, reasoning-aware safety alignment. Warning: this paper contains example data that may be offensive, harmful, or biased.
format Preprint
id arxiv_https___arxiv_org_abs_2603_17759
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Harm or Humor: A Multimodal, Multilingual Benchmark for Overt and Covert Harmful Humor
Sharshar, Ahmed
Elgendy, Hosam
Ahmed, Saad El Dine
Rohaim, Yasser
Wang, Yuxia
Computation and Language
Artificial Intelligence
Dark humor often relies on subtle cultural nuances and implicit cues that require contextual reasoning to interpret, posing safety challenges that current static benchmarks fail to capture. To address this, we introduce a novel multimodal, multilingual benchmark for detecting and understanding harmful and offensive humor. Our manually curated dataset comprises 3,000 texts and 6,000 images in English and Arabic, alongside 1,200 videos that span English, Arabic, and language-independent (universal) contexts. Unlike standard toxicity datasets, we enforce a strict annotation guideline: distinguishing Safe jokes from Harmful ones, with the latter further classified into Explicit (overt) and Implicit (Covert) categories to probe deep reasoning. We systematically evaluate state-of-the-art (SOTA) open and closed-source models across all modalities. Our findings reveal that closed-source models significantly outperform open-source ones, with a notable difference in performance between the English and Arabic languages in both, underscoring the critical need for culturally grounded, reasoning-aware safety alignment. Warning: this paper contains example data that may be offensive, harmful, or biased.
title Harm or Humor: A Multimodal, Multilingual Benchmark for Overt and Covert Harmful Humor
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2603.17759