Self-Disguise Attack: Induce the LLM to disguise itself for AIGT detection evasion

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhou, Yinghan, Wen, Juan, Peng, Wanli, Wu, Zhengxian, Zhang, Ziwei, Xue, Yiming
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915456526319616
author Zhou, Yinghan
Wen, Juan
Peng, Wanli
Wu, Zhengxian
Zhang, Ziwei
Xue, Yiming
author_facet Zhou, Yinghan
Wen, Juan
Peng, Wanli
Wu, Zhengxian
Zhang, Ziwei
Xue, Yiming
contents AI-generated text (AIGT) detection evasion aims to reduce the detection probability of AIGT, helping to identify weaknesses in detectors and enhance their effectiveness and reliability in practical applications. Although existing evasion methods perform well, they suffer from high computational costs and text quality degradation. To address these challenges, we propose Self-Disguise Attack (SDA), a novel approach that enables Large Language Models (LLM) to actively disguise its output, reducing the likelihood of detection by classifiers. The SDA comprises two main components: the adversarial feature extractor and the retrieval-based context examples optimizer. The former generates disguise features that enable LLMs to understand how to produce more human-like text. The latter retrieves the most relevant examples from an external knowledge base as in-context examples, further enhancing the self-disguise ability of LLMs and mitigating the impact of the disguise process on the diversity of the generated text. The SDA directly employs prompts containing disguise features and optimized context examples to guide the LLM in generating detection-resistant text, thereby reducing resource consumption. Experimental results demonstrate that the SDA effectively reduces the average detection accuracy of various AIGT detectors across texts generated by three different LLMs, while maintaining the quality of AIGT.
format Preprint
id arxiv_https___arxiv_org_abs_2508_15848
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Self-Disguise Attack: Induce the LLM to disguise itself for AIGT detection evasion
Zhou, Yinghan
Wen, Juan
Peng, Wanli
Wu, Zhengxian
Zhang, Ziwei
Xue, Yiming
Cryptography and Security
Computation and Language
AI-generated text (AIGT) detection evasion aims to reduce the detection probability of AIGT, helping to identify weaknesses in detectors and enhance their effectiveness and reliability in practical applications. Although existing evasion methods perform well, they suffer from high computational costs and text quality degradation. To address these challenges, we propose Self-Disguise Attack (SDA), a novel approach that enables Large Language Models (LLM) to actively disguise its output, reducing the likelihood of detection by classifiers. The SDA comprises two main components: the adversarial feature extractor and the retrieval-based context examples optimizer. The former generates disguise features that enable LLMs to understand how to produce more human-like text. The latter retrieves the most relevant examples from an external knowledge base as in-context examples, further enhancing the self-disguise ability of LLMs and mitigating the impact of the disguise process on the diversity of the generated text. The SDA directly employs prompts containing disguise features and optimized context examples to guide the LLM in generating detection-resistant text, thereby reducing resource consumption. Experimental results demonstrate that the SDA effectively reduces the average detection accuracy of various AIGT detectors across texts generated by three different LLMs, while maintaining the quality of AIGT.
title Self-Disguise Attack: Induce the LLM to disguise itself for AIGT detection evasion
topic Cryptography and Security
Computation and Language
url https://arxiv.org/abs/2508.15848