Meme Trojan: Backdoor Attacks Against Hateful Meme Detection via Cross-Modal Triggers

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Ruofei, Lin, Hongzhan, Luo, Ziyuan, Cheung, Ka Chun, See, Simon, Ma, Jing, Wan, Renjie
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908168684044288
author Wang, Ruofei
Lin, Hongzhan
Luo, Ziyuan
Cheung, Ka Chun
See, Simon
Ma, Jing
Wan, Renjie
author_facet Wang, Ruofei
Lin, Hongzhan
Luo, Ziyuan
Cheung, Ka Chun
See, Simon
Ma, Jing
Wan, Renjie
contents Hateful meme detection aims to prevent the proliferation of hateful memes on various social media platforms. Considering its impact on social environments, this paper introduces a previously ignored but significant threat to hateful meme detection: backdoor attacks. By injecting specific triggers into meme samples, backdoor attackers can manipulate the detector to output their desired outcomes. To explore this, we propose the Meme Trojan framework to initiate backdoor attacks on hateful meme detection. Meme Trojan involves creating a novel Cross-Modal Trigger (CMT) and a learnable trigger augmentor to enhance the trigger pattern according to each input sample. Due to the cross-modal property, the proposed CMT can effectively initiate backdoor attacks on hateful meme detectors under an automatic application scenario. Additionally, the injection position and size of our triggers are adaptive to the texts contained in the meme, which ensures that the trigger is seamlessly integrated with the meme content. Our approach outperforms the state-of-the-art backdoor attack methods, showing significant improvements in effectiveness and stealthiness. We believe that this paper will draw more attention to the potential threat posed by backdoor attacks on hateful meme detection.
format Preprint
id arxiv_https___arxiv_org_abs_2412_15503
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Meme Trojan: Backdoor Attacks Against Hateful Meme Detection via Cross-Modal Triggers
Wang, Ruofei
Lin, Hongzhan
Luo, Ziyuan
Cheung, Ka Chun
See, Simon
Ma, Jing
Wan, Renjie
Cryptography and Security
Hateful meme detection aims to prevent the proliferation of hateful memes on various social media platforms. Considering its impact on social environments, this paper introduces a previously ignored but significant threat to hateful meme detection: backdoor attacks. By injecting specific triggers into meme samples, backdoor attackers can manipulate the detector to output their desired outcomes. To explore this, we propose the Meme Trojan framework to initiate backdoor attacks on hateful meme detection. Meme Trojan involves creating a novel Cross-Modal Trigger (CMT) and a learnable trigger augmentor to enhance the trigger pattern according to each input sample. Due to the cross-modal property, the proposed CMT can effectively initiate backdoor attacks on hateful meme detectors under an automatic application scenario. Additionally, the injection position and size of our triggers are adaptive to the texts contained in the meme, which ensures that the trigger is seamlessly integrated with the meme content. Our approach outperforms the state-of-the-art backdoor attack methods, showing significant improvements in effectiveness and stealthiness. We believe that this paper will draw more attention to the potential threat posed by backdoor attacks on hateful meme detection.
title Meme Trojan: Backdoor Attacks Against Hateful Meme Detection via Cross-Modal Triggers
topic Cryptography and Security
url https://arxiv.org/abs/2412.15503