Stress-testing Machine Generated Text Detection: Shifting Language Models Writing Style to Fool Detectors
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866910975995674624 |
|---|---|
| author | Pedrotti, Andrea Papucci, Michele Ciaccio, Cristiano Miaschi, Alessio Puccetti, Giovanni Dell'Orletta, Felice Esuli, Andrea |
| author_facet | Pedrotti, Andrea Papucci, Michele Ciaccio, Cristiano Miaschi, Alessio Puccetti, Giovanni Dell'Orletta, Felice Esuli, Andrea |
| contents | Recent advancements in Generative AI and Large Language Models (LLMs) have enabled the creation of highly realistic synthetic content, raising concerns about the potential for malicious use, such as misinformation and manipulation. Moreover, detecting Machine-Generated Text (MGT) remains challenging due to the lack of robust benchmarks that assess generalization to real-world scenarios. In this work, we present a pipeline to test the resilience of state-of-the-art MGT detectors (e.g., Mage, Radar, LLM-DetectAIve) to linguistically informed adversarial attacks. To challenge the detectors, we fine-tune language models using Direct Preference Optimization (DPO) to shift the MGT style toward human-written text (HWT). This exploits the detectors' reliance on stylistic clues, making new generations more challenging to detect. Additionally, we analyze the linguistic shifts induced by the alignment and which features are used by detectors to detect MGT texts. Our results show that detectors can be easily fooled with relatively few examples, resulting in a significant drop in detection performance. This highlights the importance of improving detection methods and making them robust to unseen in-domain texts. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2505_24523 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Stress-testing Machine Generated Text Detection: Shifting Language Models Writing Style to Fool Detectors Pedrotti, Andrea Papucci, Michele Ciaccio, Cristiano Miaschi, Alessio Puccetti, Giovanni Dell'Orletta, Felice Esuli, Andrea Computation and Language Artificial Intelligence Recent advancements in Generative AI and Large Language Models (LLMs) have enabled the creation of highly realistic synthetic content, raising concerns about the potential for malicious use, such as misinformation and manipulation. Moreover, detecting Machine-Generated Text (MGT) remains challenging due to the lack of robust benchmarks that assess generalization to real-world scenarios. In this work, we present a pipeline to test the resilience of state-of-the-art MGT detectors (e.g., Mage, Radar, LLM-DetectAIve) to linguistically informed adversarial attacks. To challenge the detectors, we fine-tune language models using Direct Preference Optimization (DPO) to shift the MGT style toward human-written text (HWT). This exploits the detectors' reliance on stylistic clues, making new generations more challenging to detect. Additionally, we analyze the linguistic shifts induced by the alignment and which features are used by detectors to detect MGT texts. Our results show that detectors can be easily fooled with relatively few examples, resulting in a significant drop in detection performance. This highlights the importance of improving detection methods and making them robust to unseen in-domain texts. |
| title | Stress-testing Machine Generated Text Detection: Shifting Language Models Writing Style to Fool Detectors |
| topic | Computation and Language Artificial Intelligence |
| url | https://arxiv.org/abs/2505.24523 |