Stress-testing Machine Generated Text Detection: Shifting Language Models Writing Style to Fool Detectors

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Pedrotti, Andrea, Papucci, Michele, Ciaccio, Cristiano, Miaschi, Alessio, Puccetti, Giovanni, Dell'Orletta, Felice, Esuli, Andrea
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910975995674624
author Pedrotti, Andrea
Papucci, Michele
Ciaccio, Cristiano
Miaschi, Alessio
Puccetti, Giovanni
Dell'Orletta, Felice
Esuli, Andrea
author_facet Pedrotti, Andrea
Papucci, Michele
Ciaccio, Cristiano
Miaschi, Alessio
Puccetti, Giovanni
Dell'Orletta, Felice
Esuli, Andrea
contents Recent advancements in Generative AI and Large Language Models (LLMs) have enabled the creation of highly realistic synthetic content, raising concerns about the potential for malicious use, such as misinformation and manipulation. Moreover, detecting Machine-Generated Text (MGT) remains challenging due to the lack of robust benchmarks that assess generalization to real-world scenarios. In this work, we present a pipeline to test the resilience of state-of-the-art MGT detectors (e.g., Mage, Radar, LLM-DetectAIve) to linguistically informed adversarial attacks. To challenge the detectors, we fine-tune language models using Direct Preference Optimization (DPO) to shift the MGT style toward human-written text (HWT). This exploits the detectors' reliance on stylistic clues, making new generations more challenging to detect. Additionally, we analyze the linguistic shifts induced by the alignment and which features are used by detectors to detect MGT texts. Our results show that detectors can be easily fooled with relatively few examples, resulting in a significant drop in detection performance. This highlights the importance of improving detection methods and making them robust to unseen in-domain texts.
format Preprint
id arxiv_https___arxiv_org_abs_2505_24523
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Stress-testing Machine Generated Text Detection: Shifting Language Models Writing Style to Fool Detectors
Pedrotti, Andrea
Papucci, Michele
Ciaccio, Cristiano
Miaschi, Alessio
Puccetti, Giovanni
Dell'Orletta, Felice
Esuli, Andrea
Computation and Language
Artificial Intelligence
Recent advancements in Generative AI and Large Language Models (LLMs) have enabled the creation of highly realistic synthetic content, raising concerns about the potential for malicious use, such as misinformation and manipulation. Moreover, detecting Machine-Generated Text (MGT) remains challenging due to the lack of robust benchmarks that assess generalization to real-world scenarios. In this work, we present a pipeline to test the resilience of state-of-the-art MGT detectors (e.g., Mage, Radar, LLM-DetectAIve) to linguistically informed adversarial attacks. To challenge the detectors, we fine-tune language models using Direct Preference Optimization (DPO) to shift the MGT style toward human-written text (HWT). This exploits the detectors' reliance on stylistic clues, making new generations more challenging to detect. Additionally, we analyze the linguistic shifts induced by the alignment and which features are used by detectors to detect MGT texts. Our results show that detectors can be easily fooled with relatively few examples, resulting in a significant drop in detection performance. This highlights the importance of improving detection methods and making them robust to unseen in-domain texts.
title Stress-testing Machine Generated Text Detection: Shifting Language Models Writing Style to Fool Detectors
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2505.24523