"Humor, Art, or Misinformation?": A Multimodal Dataset for Intent-Aware Synthetic Image Detection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Skoularikis, Anastasios, Papadopoulos, Stefanos-Iordanis, Papadopoulos, Symeon, Petrantonakis, Panagiotis C.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908526694105088
author Skoularikis, Anastasios
Papadopoulos, Stefanos-Iordanis
Papadopoulos, Symeon
Petrantonakis, Panagiotis C.
author_facet Skoularikis, Anastasios
Papadopoulos, Stefanos-Iordanis
Papadopoulos, Symeon
Petrantonakis, Panagiotis C.
contents Recent advances in multimodal AI have enabled progress in detecting synthetic and out-of-context content. However, existing efforts largely overlook the intent behind AI-generated images. To fill this gap, we introduce S-HArM, a multimodal dataset for intent-aware classification, comprising 9,576 "in the wild" image-text pairs from Twitter/X and Reddit, labeled as Humor/Satire, Art, or Misinformation. Additionally, we explore three prompting strategies (image-guided, description-guided, and multimodally-guided) to construct a large-scale synthetic training dataset with Stable Diffusion. We conduct an extensive comparative study including modality fusion, contrastive learning, reconstruction networks, attention mechanisms, and large vision-language models. Our results show that models trained on image- and multimodally-guided data generalize better to "in the wild" content, due to preserved visual context. However, overall performance remains limited, highlighting the complexity of inferring intent and the need for specialized architectures.
format Preprint
id arxiv_https___arxiv_org_abs_2508_20670
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle "Humor, Art, or Misinformation?": A Multimodal Dataset for Intent-Aware Synthetic Image Detection
Skoularikis, Anastasios
Papadopoulos, Stefanos-Iordanis
Papadopoulos, Symeon
Petrantonakis, Panagiotis C.
Computer Vision and Pattern Recognition
Multimedia
Recent advances in multimodal AI have enabled progress in detecting synthetic and out-of-context content. However, existing efforts largely overlook the intent behind AI-generated images. To fill this gap, we introduce S-HArM, a multimodal dataset for intent-aware classification, comprising 9,576 "in the wild" image-text pairs from Twitter/X and Reddit, labeled as Humor/Satire, Art, or Misinformation. Additionally, we explore three prompting strategies (image-guided, description-guided, and multimodally-guided) to construct a large-scale synthetic training dataset with Stable Diffusion. We conduct an extensive comparative study including modality fusion, contrastive learning, reconstruction networks, attention mechanisms, and large vision-language models. Our results show that models trained on image- and multimodally-guided data generalize better to "in the wild" content, due to preserved visual context. However, overall performance remains limited, highlighting the complexity of inferring intent and the need for specialized architectures.
title "Humor, Art, or Misinformation?": A Multimodal Dataset for Intent-Aware Synthetic Image Detection
topic Computer Vision and Pattern Recognition
Multimedia
url https://arxiv.org/abs/2508.20670