Your Language Model Can Secretly Write Like Humans: Contrastive Paraphrase Attacks on LLM-Generated Text Detectors

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Fang, Hao, Kong, Jiawei, Zhuang, Tianqu, Qiu, Yixiang, Gao, Kuofeng, Chen, Bin, Xia, Shu-Tao, Wang, Yaowei, Zhang, Min
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916943905161216
author Fang, Hao
Kong, Jiawei
Zhuang, Tianqu
Qiu, Yixiang
Gao, Kuofeng
Chen, Bin
Xia, Shu-Tao
Wang, Yaowei
Zhang, Min
author_facet Fang, Hao
Kong, Jiawei
Zhuang, Tianqu
Qiu, Yixiang
Gao, Kuofeng
Chen, Bin
Xia, Shu-Tao
Wang, Yaowei
Zhang, Min
contents The misuse of large language models (LLMs), such as academic plagiarism, has driven the development of detectors to identify LLM-generated texts. To bypass these detectors, paraphrase attacks have emerged to purposely rewrite these texts to evade detection. Despite the success, existing methods require substantial data and computational budgets to train a specialized paraphraser, and their attack efficacy greatly reduces when faced with advanced detection algorithms. To address this, we propose \textbf{Co}ntrastive \textbf{P}araphrase \textbf{A}ttack (CoPA), a training-free method that effectively deceives text detectors using off-the-shelf LLMs. The first step is to carefully craft instructions that encourage LLMs to produce more human-like texts. Nonetheless, we observe that the inherent statistical biases of LLMs can still result in some generated texts carrying certain machine-like attributes that can be captured by detectors. To overcome this, CoPA constructs an auxiliary machine-like word distribution as a contrast to the human-like distribution generated by the LLM. By subtracting the machine-like patterns from the human-like distribution during the decoding process, CoPA is able to produce sentences that are less discernible by text detectors. Our theoretical analysis suggests the superiority of the proposed attack. Extensive experiments validate the effectiveness of CoPA in fooling text detectors across various scenarios.
format Preprint
id arxiv_https___arxiv_org_abs_2505_15337
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Your Language Model Can Secretly Write Like Humans: Contrastive Paraphrase Attacks on LLM-Generated Text Detectors
Fang, Hao
Kong, Jiawei
Zhuang, Tianqu
Qiu, Yixiang
Gao, Kuofeng
Chen, Bin
Xia, Shu-Tao
Wang, Yaowei
Zhang, Min
Computation and Language
Artificial Intelligence
The misuse of large language models (LLMs), such as academic plagiarism, has driven the development of detectors to identify LLM-generated texts. To bypass these detectors, paraphrase attacks have emerged to purposely rewrite these texts to evade detection. Despite the success, existing methods require substantial data and computational budgets to train a specialized paraphraser, and their attack efficacy greatly reduces when faced with advanced detection algorithms. To address this, we propose \textbf{Co}ntrastive \textbf{P}araphrase \textbf{A}ttack (CoPA), a training-free method that effectively deceives text detectors using off-the-shelf LLMs. The first step is to carefully craft instructions that encourage LLMs to produce more human-like texts. Nonetheless, we observe that the inherent statistical biases of LLMs can still result in some generated texts carrying certain machine-like attributes that can be captured by detectors. To overcome this, CoPA constructs an auxiliary machine-like word distribution as a contrast to the human-like distribution generated by the LLM. By subtracting the machine-like patterns from the human-like distribution during the decoding process, CoPA is able to produce sentences that are less discernible by text detectors. Our theoretical analysis suggests the superiority of the proposed attack. Extensive experiments validate the effectiveness of CoPA in fooling text detectors across various scenarios.
title Your Language Model Can Secretly Write Like Humans: Contrastive Paraphrase Attacks on LLM-Generated Text Detectors
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2505.15337