Stumbling Blocks: Stress Testing the Robustness of Machine-Generated Text Detectors Under Attacks

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Wang, Yichen, Feng, Shangbin, Hou, Abe Bohan, Pu, Xiao, Shen, Chao, Liu, Xiaoming, Tsvetkov, Yulia, He, Tianxing
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866913236971945984
author Wang, Yichen
Feng, Shangbin
Hou, Abe Bohan
Pu, Xiao
Shen, Chao
Liu, Xiaoming
Tsvetkov, Yulia
He, Tianxing
author_facet Wang, Yichen
Feng, Shangbin
Hou, Abe Bohan
Pu, Xiao
Shen, Chao
Liu, Xiaoming
Tsvetkov, Yulia
He, Tianxing
contents The widespread use of large language models (LLMs) is increasing the demand for methods that detect machine-generated text to prevent misuse. The goal of our study is to stress test the detectors' robustness to malicious attacks under realistic scenarios. We comprehensively study the robustness of popular machine-generated text detectors under attacks from diverse categories: editing, paraphrasing, prompting, and co-generating. Our attacks assume limited access to the generator LLMs, and we compare the performance of detectors on different attacks under different budget levels. Our experiments reveal that almost none of the existing detectors remain robust under all the attacks, and all detectors exhibit different loopholes. Averaging all detectors, the performance drops by 35% across all attacks. Further, we investigate the reasons behind these defects and propose initial out-of-the-box patches to improve robustness.
format Preprint
id arxiv_https___arxiv_org_abs_2402_11638
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Stumbling Blocks: Stress Testing the Robustness of Machine-Generated Text Detectors Under Attacks
Wang, Yichen
Feng, Shangbin
Hou, Abe Bohan
Pu, Xiao
Shen, Chao
Liu, Xiaoming
Tsvetkov, Yulia
He, Tianxing
Computation and Language
The widespread use of large language models (LLMs) is increasing the demand for methods that detect machine-generated text to prevent misuse. The goal of our study is to stress test the detectors' robustness to malicious attacks under realistic scenarios. We comprehensively study the robustness of popular machine-generated text detectors under attacks from diverse categories: editing, paraphrasing, prompting, and co-generating. Our attacks assume limited access to the generator LLMs, and we compare the performance of detectors on different attacks under different budget levels. Our experiments reveal that almost none of the existing detectors remain robust under all the attacks, and all detectors exhibit different loopholes. Averaging all detectors, the performance drops by 35% across all attacks. Further, we investigate the reasons behind these defects and propose initial out-of-the-box patches to improve robustness.
title Stumbling Blocks: Stress Testing the Robustness of Machine-Generated Text Detectors Under Attacks
topic Computation and Language
url https://arxiv.org/abs/2402.11638