Fooling the Watchers: Breaking AIGC Detectors via Semantic Prompt Attacks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hao, Run, Ying, Peng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910973655252992
author Hao, Run
Ying, Peng
author_facet Hao, Run
Ying, Peng
contents The rise of text-to-image (T2I) models has enabled the synthesis of photorealistic human portraits, raising serious concerns about identity misuse and the robustness of AIGC detectors. In this work, we propose an automated adversarial prompt generation framework that leverages a grammar tree structure and a variant of the Monte Carlo tree search algorithm to systematically explore the semantic prompt space. Our method generates diverse, controllable prompts that consistently evade both open-source and commercial AIGC detectors. Extensive experiments across multiple T2I models validate its effectiveness, and the approach ranked first in a real-world adversarial AIGC detection competition. Beyond attack scenarios, our method can also be used to construct high-quality adversarial datasets, providing valuable resources for training and evaluating more robust AIGC detection and defense systems.
format Preprint
id arxiv_https___arxiv_org_abs_2505_23192
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Fooling the Watchers: Breaking AIGC Detectors via Semantic Prompt Attacks
Hao, Run
Ying, Peng
Computer Vision and Pattern Recognition
Artificial Intelligence
Cryptography and Security
The rise of text-to-image (T2I) models has enabled the synthesis of photorealistic human portraits, raising serious concerns about identity misuse and the robustness of AIGC detectors. In this work, we propose an automated adversarial prompt generation framework that leverages a grammar tree structure and a variant of the Monte Carlo tree search algorithm to systematically explore the semantic prompt space. Our method generates diverse, controllable prompts that consistently evade both open-source and commercial AIGC detectors. Extensive experiments across multiple T2I models validate its effectiveness, and the approach ranked first in a real-world adversarial AIGC detection competition. Beyond attack scenarios, our method can also be used to construct high-quality adversarial datasets, providing valuable resources for training and evaluating more robust AIGC detection and defense systems.
title Fooling the Watchers: Breaking AIGC Detectors via Semantic Prompt Attacks
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Cryptography and Security
url https://arxiv.org/abs/2505.23192