Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2603.17759 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866910059949195264 |
|---|---|
| author | Sharshar, Ahmed Elgendy, Hosam Ahmed, Saad El Dine Rohaim, Yasser Wang, Yuxia |
| author_facet | Sharshar, Ahmed Elgendy, Hosam Ahmed, Saad El Dine Rohaim, Yasser Wang, Yuxia |
| contents | Dark humor often relies on subtle cultural nuances and implicit cues that require contextual reasoning to interpret, posing safety challenges that current static benchmarks fail to capture. To address this, we introduce a novel multimodal, multilingual benchmark for detecting and understanding harmful and offensive humor. Our manually curated dataset comprises 3,000 texts and 6,000 images in English and Arabic, alongside 1,200 videos that span English, Arabic, and language-independent (universal) contexts. Unlike standard toxicity datasets, we enforce a strict annotation guideline: distinguishing Safe jokes from Harmful ones, with the latter further classified into Explicit (overt) and Implicit (Covert) categories to probe deep reasoning. We systematically evaluate state-of-the-art (SOTA) open and closed-source models across all modalities. Our findings reveal that closed-source models significantly outperform open-source ones, with a notable difference in performance between the English and Arabic languages in both, underscoring the critical need for culturally grounded, reasoning-aware safety alignment. Warning: this paper contains example data that may be offensive, harmful, or biased. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2603_17759 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Harm or Humor: A Multimodal, Multilingual Benchmark for Overt and Covert Harmful Humor Sharshar, Ahmed Elgendy, Hosam Ahmed, Saad El Dine Rohaim, Yasser Wang, Yuxia Computation and Language Artificial Intelligence Dark humor often relies on subtle cultural nuances and implicit cues that require contextual reasoning to interpret, posing safety challenges that current static benchmarks fail to capture. To address this, we introduce a novel multimodal, multilingual benchmark for detecting and understanding harmful and offensive humor. Our manually curated dataset comprises 3,000 texts and 6,000 images in English and Arabic, alongside 1,200 videos that span English, Arabic, and language-independent (universal) contexts. Unlike standard toxicity datasets, we enforce a strict annotation guideline: distinguishing Safe jokes from Harmful ones, with the latter further classified into Explicit (overt) and Implicit (Covert) categories to probe deep reasoning. We systematically evaluate state-of-the-art (SOTA) open and closed-source models across all modalities. Our findings reveal that closed-source models significantly outperform open-source ones, with a notable difference in performance between the English and Arabic languages in both, underscoring the critical need for culturally grounded, reasoning-aware safety alignment. Warning: this paper contains example data that may be offensive, harmful, or biased. |
| title | Harm or Humor: A Multimodal, Multilingual Benchmark for Overt and Covert Harmful Humor |
| topic | Computation and Language Artificial Intelligence |
| url | https://arxiv.org/abs/2603.17759 |